Nature News Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
arXiv Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
arXiv Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
arXiv The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
arXiv The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
arXiv The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
arXiv The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
arXiv ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
arXiv The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
arXiv The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
arXiv WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
arXiv The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
arXiv The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
arXiv The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
arXiv Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
arXiv This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
arXiv The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
arXiv The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
arXiv Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
arXiv Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.