CSIAM Transactions on Applied Mathematics The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
bioRxiv The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
MIT Technology Review An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
Microsoft Research Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Terence Tao blog RSS A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
IEEE Spectrum IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
MIT Technology Review Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
arXiv The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
arXiv The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
arXiv Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
arXiv The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
arXiv The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
arXiv EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
arXiv The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
arXiv The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
arXiv TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
arXiv The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
arXiv EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
arXiv NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
arXiv Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.