Only content delivered through the publication boundary on this date is included.
Anthropic Anthropic's Frontier Red Team evaluated GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), using automated benchmarks and human-in-the-loop workflows, finding that it can autonomously build end-to-end cyber exploits (50 of 410 attempts on ExploitBench and full control-flow hijacks in 4% of trials on an internal binary exploitation benchmark) and that simple techniques bypassed its safeguards in 64%–100% of simulated tests, while those techniques did not succeed against safeguarded Claude models in their testing.
Anthropic's Frontier Red Team evaluated GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), using automated benchmarks and human-in-the-loop workflows, finding that it can autonomously build end-to-end cyber exploits (50 of 410 attempts on ExploitBench and full control-flow hijacks in 4% of trials on an internal binary exploitation benchmark) and that simple techniques bypassed its safeguards in 64%–100% of simulated tests, while those techniques did not succeed against safeguarded Claude models in their testing.
Anthropic's Frontier Red Team evaluated GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), using automated benchmarks and human-in-the-loop workflows, finding that it can autonomously build end-to-end cyber exploits (50 of 410 attempts on ExploitBench and full control-flow hijacks in 4% of trials on an internal binary exploitation benchmark) and that simple techniques bypassed its safeguards in 64%–100% of simulated tests, while those techniques did not succeed against safeguarded Claude models in their testing.
Anthropic's Frontier Red Team evaluated GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), using automated benchmarks and human-in-the-loop workflows, finding that it can autonomously build end-to-end cyber exploits (50 of 410 attempts on ExploitBench and full control-flow hijacks in 4% of trials on an internal binary exploitation benchmark) and that simple techniques bypassed its safeguards in 64%–100% of simulated tests, while those techniques did not succeed against safeguarded Claude models in their testing.
Microsoft Research Microsoft Research introduces Quine, a research system combining a world model of biology trained jointly across modalities including sequence, structure, function, cellular state, and imaging with a harness connecting scientific tools, literature, the wet lab, and researchers; with the Broad Institute it used the system in pancreatic ductal adenocarcinoma (PDAC) to predict and prioritize thousands of compounds by their potential to shift tumor cells between therapeutically relevant states, wet-lab assays showed Quine's highest-ranked compounds produced the largest intended shifts, the whole process from narrowing the search space to prioritizing a handful of candidates took just one weekend, and the experiments also bore out the model's prediction of a distinct third phenotype.
Microsoft Research introduces Quine, a research system combining a world model of biology trained jointly across modalities including sequence, structure, function, cellular state, and imaging with a harness connecting scientific tools, literature, the wet lab, and researchers; with the Broad Institute it used the system in pancreatic ductal adenocarcinoma (PDAC) to predict and prioritize thousands of compounds by their potential to shift tumor cells between therapeutically relevant states, wet-lab assays showed Quine's highest-ranked compounds produced the largest intended shifts, the whole process from narrowing the search space to prioritizing a handful of candidates took just one weekend, and the experiments also bore out the model's prediction of a distinct third phenotype.
Microsoft Research introduces Quine, a research system combining a world model of biology trained jointly across modalities including sequence, structure, function, cellular state, and imaging with a harness connecting scientific tools, literature, the wet lab, and researchers; with the Broad Institute it used the system in pancreatic ductal adenocarcinoma (PDAC) to predict and prioritize thousands of compounds by their potential to shift tumor cells between therapeutically relevant states, wet-lab assays showed Quine's highest-ranked compounds produced the largest intended shifts, the whole process from narrowing the search space to prioritizing a handful of candidates took just one weekend, and the experiments also bore out the model's prediction of a distinct third phenotype.
Microsoft Research introduces Quine, a research system combining a world model of biology trained jointly across modalities including sequence, structure, function, cellular state, and imaging with a harness connecting scientific tools, literature, the wet lab, and researchers; with the Broad Institute it used the system in pancreatic ductal adenocarcinoma (PDAC) to predict and prioritize thousands of compounds by their potential to shift tumor cells between therapeutically relevant states, wet-lab assays showed Quine's highest-ranked compounds produced the largest intended shifts, the whole process from narrowing the search space to prioritizing a handful of candidates took just one weekend, and the experiments also bore out the model's prediction of a distinct third phenotype.
IEEE Spectrum This article traces how animal-testing alternatives known as NAMs—organ chips, organoids and computational simulations—moved from a lung-on-a-chip paper that Science asked to be backed up with mouse experiments, to the 2022 FDA Modernization Act 2.0 authorizing NAMs in preclinical studies, a 2025 FDA pledge to make animal studies the exception, and a September 2026 rule that would replace "animal tests" with "nonclinical tests" in drug regulations, while identifying validation standardization, the high cost of head-to-head comparisons and research-culture inertia as the main bottlenecks.
This article traces how animal-testing alternatives known as NAMs—organ chips, organoids and computational simulations—moved from a lung-on-a-chip paper that Science asked to be backed up with mouse experiments, to the 2022 FDA Modernization Act 2.0 authorizing NAMs in preclinical studies, a 2025 FDA pledge to make animal studies the exception, and a September 2026 rule that would replace "animal tests" with "nonclinical tests" in drug regulations, while identifying validation standardization, the high cost of head-to-head comparisons and research-culture inertia as the main bottlenecks.
This article traces how animal-testing alternatives known as NAMs—organ chips, organoids and computational simulations—moved from a lung-on-a-chip paper that Science asked to be backed up with mouse experiments, to the 2022 FDA Modernization Act 2.0 authorizing NAMs in preclinical studies, a 2025 FDA pledge to make animal studies the exception, and a September 2026 rule that would replace "animal tests" with "nonclinical tests" in drug regulations, while identifying validation standardization, the high cost of head-to-head comparisons and research-culture inertia as the main bottlenecks.
This article traces how animal-testing alternatives known as NAMs—organ chips, organoids and computational simulations—moved from a lung-on-a-chip paper that Science asked to be backed up with mouse experiments, to the 2022 FDA Modernization Act 2.0 authorizing NAMs in preclinical studies, a 2025 FDA pledge to make animal studies the exception, and a September 2026 rule that would replace "animal tests" with "nonclinical tests" in drug regulations, while identifying validation standardization, the high cost of head-to-head comparisons and research-culture inertia as the main bottlenecks.
MIT Technology Review MIT Technology Review's The Download newsletter reports that Anthropic announced its new molecular biology lab's first discovery: its AI agents flagged a previously uncatalogued pattern surrounding an enzyme, a pattern "reminiscent" of what led to the gene-editing technology CRISPR, but biologists pushed back, with some questioning whether merely finding the pattern amounted to a discovery and another saying his team had already discovered the same pattern, raising questions about whether Anthropic's system had learned from his conversations with Claude.
MIT Technology Review's The Download newsletter reports that Anthropic announced its new molecular biology lab's first discovery: its AI agents flagged a previously uncatalogued pattern surrounding an enzyme, a pattern "reminiscent" of what led to the gene-editing technology CRISPR, but biologists pushed back, with some questioning whether merely finding the pattern amounted to a discovery and another saying his team had already discovered the same pattern, raising questions about whether Anthropic's system had learned from his conversations with Claude.
MIT Technology Review's The Download newsletter reports that Anthropic announced its new molecular biology lab's first discovery: its AI agents flagged a previously uncatalogued pattern surrounding an enzyme, a pattern "reminiscent" of what led to the gene-editing technology CRISPR, but biologists pushed back, with some questioning whether merely finding the pattern amounted to a discovery and another saying his team had already discovered the same pattern, raising questions about whether Anthropic's system had learned from his conversations with Claude.
MIT Technology Review's The Download newsletter reports that Anthropic announced its new molecular biology lab's first discovery: its AI agents flagged a previously uncatalogued pattern surrounding an enzyme, a pattern "reminiscent" of what led to the gene-editing technology CRISPR, but biologists pushed back, with some questioning whether merely finding the pattern amounted to a discovery and another saying his team had already discovered the same pattern, raising questions about whether Anthropic's system had learned from his conversations with Claude.
World Health Organization On 22 September 2026, WHO and Knowledge Ecology International (KEI) co-organized an expert panel, "Intellectual Property and Traditional Medicine: Rethinking Access, Benefit Sharing and Data Governance," on the sidelines of the 53rd Inter Government Consultations at WIPO, where participants noted that AI and digital technologies can accelerate drug discovery based on traditional medicine knowledge while making contributions by knowledge holders harder to trace, and put forward concrete proposals including federated data governance, prior informed consent and tailored benefit-sharing arrangements, while stressing that Indigenous Peoples and local communities should be recognized as partners in stewardship and governance rather than merely as sources of knowledge.
On 22 September 2026, WHO and Knowledge Ecology International (KEI) co-organized an expert panel, "Intellectual Property and Traditional Medicine: Rethinking Access, Benefit Sharing and Data Governance," on the sidelines of the 53rd Inter Government Consultations at WIPO, where participants noted that AI and digital technologies can accelerate drug discovery based on traditional medicine knowledge while making contributions by knowledge holders harder to trace, and put forward concrete proposals including federated data governance, prior informed consent and tailored benefit-sharing arrangements, while stressing that Indigenous Peoples and local communities should be recognized as partners in stewardship and governance rather than merely as sources of knowledge.
On 22 September 2026, WHO and Knowledge Ecology International (KEI) co-organized an expert panel, "Intellectual Property and Traditional Medicine: Rethinking Access, Benefit Sharing and Data Governance," on the sidelines of the 53rd Inter Government Consultations at WIPO, where participants noted that AI and digital technologies can accelerate drug discovery based on traditional medicine knowledge while making contributions by knowledge holders harder to trace, and put forward concrete proposals including federated data governance, prior informed consent and tailored benefit-sharing arrangements, while stressing that Indigenous Peoples and local communities should be recognized as partners in stewardship and governance rather than merely as sources of knowledge.
On 22 September 2026, WHO and Knowledge Ecology International (KEI) co-organized an expert panel, "Intellectual Property and Traditional Medicine: Rethinking Access, Benefit Sharing and Data Governance," on the sidelines of the 53rd Inter Government Consultations at WIPO, where participants noted that AI and digital technologies can accelerate drug discovery based on traditional medicine knowledge while making contributions by knowledge holders harder to trace, and put forward concrete proposals including federated data governance, prior informed consent and tailored benefit-sharing arrangements, while stressing that Indigenous Peoples and local communities should be recognized as partners in stewardship and governance rather than merely as sources of knowledge.
MIT Technology Review This sponsored article, provided by HPE and not written by MIT Technology Review's editorial staff, argues that as AI moves from isolated pilots into production portfolios (assistants, retrieval-and-knowledge systems, agentic applications), a consumption-only approach turns AI spending into a hard-to-forecast variable monthly line item, so enterprises should assess workload by workload the 'crossover point' at which sustained use makes owning and operating capacity potentially more economical than buying one request at a time, while stressing that the capital decision is only half the equation and that adoption, governance, and continued expansion of high-value use cases are needed to keep that capacity productive.
This sponsored article, provided by HPE and not written by MIT Technology Review's editorial staff, argues that as AI moves from isolated pilots into production portfolios (assistants, retrieval-and-knowledge systems, agentic applications), a consumption-only approach turns AI spending into a hard-to-forecast variable monthly line item, so enterprises should assess workload by workload the 'crossover point' at which sustained use makes owning and operating capacity potentially more economical than buying one request at a time, while stressing that the capital decision is only half the equation and that adoption, governance, and continued expansion of high-value use cases are needed to keep that capacity productive.
This sponsored article, provided by HPE and not written by MIT Technology Review's editorial staff, argues that as AI moves from isolated pilots into production portfolios (assistants, retrieval-and-knowledge systems, agentic applications), a consumption-only approach turns AI spending into a hard-to-forecast variable monthly line item, so enterprises should assess workload by workload the 'crossover point' at which sustained use makes owning and operating capacity potentially more economical than buying one request at a time, while stressing that the capital decision is only half the equation and that adoption, governance, and continued expansion of high-value use cases are needed to keep that capacity productive.
This sponsored article, provided by HPE and not written by MIT Technology Review's editorial staff, argues that as AI moves from isolated pilots into production portfolios (assistants, retrieval-and-knowledge systems, agentic applications), a consumption-only approach turns AI spending into a hard-to-forecast variable monthly line item, so enterprises should assess workload by workload the 'crossover point' at which sustained use makes owning and operating capacity potentially more economical than buying one request at a time, while stressing that the capital decision is only half the equation and that adoption, governance, and continued expansion of high-value use cases are needed to keep that capacity productive.
OpenAI News The announcement introduces GPT-6.1 Sol, described as offering near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices.
The announcement introduces GPT-6.1 Sol, described as offering near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices.
The announcement introduces GPT-6.1 Sol, described as offering near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices.
The announcement introduces GPT-6.1 Sol, described as offering near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices.
OpenAI News At DevDay 2026, OpenAI announced more than 20 items spanning GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
At DevDay 2026, OpenAI announced more than 20 items spanning GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
At DevDay 2026, OpenAI announced more than 20 items spanning GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
At DevDay 2026, OpenAI announced more than 20 items spanning GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
MIT News - Artificial intelligence In Philosophical Perspectives, MIT's Brian Hedden and Manish Raghavan systematically evaluate major objections to algorithmic monoculture—where one algorithm makes all decisions in a domain—arguing that objections such as systematic exclusion do not hold, proving mathematically that monoculture tends to create information echo chambers that hinder exploration, and showing through a series of hiring simulations that bundling algorithms into an "ensemble" can sometimes overcome this limitation so that monoculture performs as well as or better than polyculture.
In Philosophical Perspectives, MIT's Brian Hedden and Manish Raghavan systematically evaluate major objections to algorithmic monoculture—where one algorithm makes all decisions in a domain—arguing that objections such as systematic exclusion do not hold, proving mathematically that monoculture tends to create information echo chambers that hinder exploration, and showing through a series of hiring simulations that bundling algorithms into an "ensemble" can sometimes overcome this limitation so that monoculture performs as well as or better than polyculture.
In Philosophical Perspectives, MIT's Brian Hedden and Manish Raghavan systematically evaluate major objections to algorithmic monoculture—where one algorithm makes all decisions in a domain—arguing that objections such as systematic exclusion do not hold, proving mathematically that monoculture tends to create information echo chambers that hinder exploration, and showing through a series of hiring simulations that bundling algorithms into an "ensemble" can sometimes overcome this limitation so that monoculture performs as well as or better than polyculture.
In Philosophical Perspectives, MIT's Brian Hedden and Manish Raghavan systematically evaluate major objections to algorithmic monoculture—where one algorithm makes all decisions in a domain—arguing that objections such as systematic exclusion do not hold, proving mathematically that monoculture tends to create information echo chambers that hinder exploration, and showing through a series of hiring simulations that bundling algorithms into an "ensemble" can sometimes overcome this limitation so that monoculture performs as well as or better than polyculture.
MIT News - Artificial intelligence In her new book 'Artificial Intimacy: Who We Become When We Talk to Machines,' MIT professor of the social studies of science and technology Sherry Turkle draws on surveyed evidence and new interview research to examine chatbots across the stages of human life, concluding that chatbot use, though often felt as a short-term salve, is broadly detrimental to human development and social connectivity, and calling for a pushback movement akin to those against phones in schools and youth social media use.
In her new book 'Artificial Intimacy: Who We Become When We Talk to Machines,' MIT professor of the social studies of science and technology Sherry Turkle draws on surveyed evidence and new interview research to examine chatbots across the stages of human life, concluding that chatbot use, though often felt as a short-term salve, is broadly detrimental to human development and social connectivity, and calling for a pushback movement akin to those against phones in schools and youth social media use.
In her new book 'Artificial Intimacy: Who We Become When We Talk to Machines,' MIT professor of the social studies of science and technology Sherry Turkle draws on surveyed evidence and new interview research to examine chatbots across the stages of human life, concluding that chatbot use, though often felt as a short-term salve, is broadly detrimental to human development and social connectivity, and calling for a pushback movement akin to those against phones in schools and youth social media use.
In her new book 'Artificial Intimacy: Who We Become When We Talk to Machines,' MIT professor of the social studies of science and technology Sherry Turkle draws on surveyed evidence and new interview research to examine chatbots across the stages of human life, concluding that chatbot use, though often felt as a short-term salve, is broadly detrimental to human development and social connectivity, and calling for a pushback movement akin to those against phones in schools and youth social media use.
bioRxiv The study benchmarked a doxycycline-inducible, hTERT-immortalized iPSC-derived mesenchymal stromal cell line engineered to overexpress TGFB1 (TGFB1-iMSCs) against multiple adipose tissue-derived MSC (MSC(AT)) donors across predefined immunomodulatory and angiogenic potency attributes, finding that TGFB1-iMSCs were smaller and more circular with comparable or higher proliferative rates, had a distinct angiogenic signature (EDIL3, EDN1, PDGFA) and nine differentially expressed microRNAs, secreted less VEGF with intermediate HUVEC tube formation, yet matched or exceeded all MSC(AT) donors in a monocyte-macrophage transwell immunomodulatory assay; in a murine DMM post-traumatic osteoarthritis model, a single intra-articular injection of TGFB1-iMSCs, but not MSC(AT), reduced total synovial macr
The study benchmarked a doxycycline-inducible, hTERT-immortalized iPSC-derived mesenchymal stromal cell line engineered to overexpress TGFB1 (TGFB1-iMSCs) against multiple adipose tissue-derived MSC (MSC(AT)) donors across predefined immunomodulatory and angiogenic potency attributes, finding that TGFB1-iMSCs were smaller and more circular with comparable or higher proliferative rates, had a distinct angiogenic signature (EDIL3, EDN1, PDGFA) and nine differentially expressed microRNAs, secreted less VEGF with intermediate HUVEC tube formation, yet matched or exceeded all MSC(AT) donors in a monocyte-macrophage transwell immunomodulatory assay; in a murine DMM post-traumatic osteoarthritis model, a single intra-articular injection of TGFB1-iMSCs, but not MSC(AT), reduced total synovial macr
The study benchmarked a doxycycline-inducible, hTERT-immortalized iPSC-derived mesenchymal stromal cell line engineered to overexpress TGFB1 (TGFB1-iMSCs) against multiple adipose tissue-derived MSC (MSC(AT)) donors across predefined immunomodulatory and angiogenic potency attributes, finding that TGFB1-iMSCs were smaller and more circular with comparable or higher proliferative rates, had a distinct angiogenic signature (EDIL3, EDN1, PDGFA) and nine differentially expressed microRNAs, secreted less VEGF with intermediate HUVEC tube formation, yet matched or exceeded all MSC(AT) donors in a monocyte-macrophage transwell immunomodulatory assay; in a murine DMM post-traumatic osteoarthritis model, a single intra-articular injection of TGFB1-iMSCs, but not MSC(AT), reduced total synovial macr
The study benchmarked a doxycycline-inducible, hTERT-immortalized iPSC-derived mesenchymal stromal cell line engineered to overexpress TGFB1 (TGFB1-iMSCs) against multiple adipose tissue-derived MSC (MSC(AT)) donors across predefined immunomodulatory and angiogenic potency attributes, finding that TGFB1-iMSCs were smaller and more circular with comparable or higher proliferative rates, had a distinct angiogenic signature (EDIL3, EDN1, PDGFA) and nine differentially expressed microRNAs, secreted less VEGF with intermediate HUVEC tube formation, yet matched or exceeded all MSC(AT) donors in a monocyte-macrophage transwell immunomodulatory assay; in a murine DMM post-traumatic osteoarthritis model, a single intra-articular injection of TGFB1-iMSCs, but not MSC(AT), reduced total synovial macr
Microsystems & Nanoengineering Researchers developed a flexible wireless wearable stethoscope based on a 25-element circular aluminum nitride (AlN) piezoelectric micromachined ultrasonic transducer (PMUT) array that achieves a packaged sensitivity of −167.5 dB, an operating bandwidth of 10 Hz–10 kHz, and a frequency-response flatness of ±0.5 dB; across multiple participants it acquired heart sounds at five standard auscultation sites with temporal correspondence to reference ECG and chest-motion signals, tracked heart rate continuously during dynamic activities such as walking and stair climbing, and, coupled with a residual neural network, classified five respiratory states (awake, asleep, apnea, rhonchi, and wheeze) with 98.7% accuracy.
Researchers developed a flexible wireless wearable stethoscope based on a 25-element circular aluminum nitride (AlN) piezoelectric micromachined ultrasonic transducer (PMUT) array that achieves a packaged sensitivity of −167.5 dB, an operating bandwidth of 10 Hz–10 kHz, and a frequency-response flatness of ±0.5 dB; across multiple participants it acquired heart sounds at five standard auscultation sites with temporal correspondence to reference ECG and chest-motion signals, tracked heart rate continuously during dynamic activities such as walking and stair climbing, and, coupled with a residual neural network, classified five respiratory states (awake, asleep, apnea, rhonchi, and wheeze) with 98.7% accuracy.
Researchers developed a flexible wireless wearable stethoscope based on a 25-element circular aluminum nitride (AlN) piezoelectric micromachined ultrasonic transducer (PMUT) array that achieves a packaged sensitivity of −167.5 dB, an operating bandwidth of 10 Hz–10 kHz, and a frequency-response flatness of ±0.5 dB; across multiple participants it acquired heart sounds at five standard auscultation sites with temporal correspondence to reference ECG and chest-motion signals, tracked heart rate continuously during dynamic activities such as walking and stair climbing, and, coupled with a residual neural network, classified five respiratory states (awake, asleep, apnea, rhonchi, and wheeze) with 98.7% accuracy.
Researchers developed a flexible wireless wearable stethoscope based on a 25-element circular aluminum nitride (AlN) piezoelectric micromachined ultrasonic transducer (PMUT) array that achieves a packaged sensitivity of −167.5 dB, an operating bandwidth of 10 Hz–10 kHz, and a frequency-response flatness of ±0.5 dB; across multiple participants it acquired heart sounds at five standard auscultation sites with temporal correspondence to reference ECG and chest-motion signals, tracked heart rate continuously during dynamic activities such as walking and stair climbing, and, coupled with a residual neural network, classified five respiratory states (awake, asleep, apnea, rhonchi, and wheeze) with 98.7% accuracy.
发表出处待核验 The study introduces PixelConfig, a differential-analysis framework that reverse-engineers Meta Pixel configurations through code-patching replays and developer-account controlled experiments, and uses Internet Archive's Wayback Machine to longitudinally compare configurations on 18K health websites against a top-10K control group from 2017 to 2024, finding that default-enabled tracking features such as automatic events and first-party cookies reached adoption rates up to 98.4%, that health websites show tracking of potentially sensitive information tied to booking medical appointments and button clicks associated with specific conditions such as erectile dysfunction, and that restriction features like Core Setup were configured on 34.3% of health websites versus 8.
The study introduces PixelConfig, a differential-analysis framework that reverse-engineers Meta Pixel configurations through code-patching replays and developer-account controlled experiments, and uses Internet Archive's Wayback Machine to longitudinally compare configurations on 18K health websites against a top-10K control group from 2017 to 2024, finding that default-enabled tracking features such as automatic events and first-party cookies reached adoption rates up to 98.4%, that health websites show tracking of potentially sensitive information tied to booking medical appointments and button clicks associated with specific conditions such as erectile dysfunction, and that restriction features like Core Setup were configured on 34.3% of health websites versus 8.
The study introduces PixelConfig, a differential-analysis framework that reverse-engineers Meta Pixel configurations through code-patching replays and developer-account controlled experiments, and uses Internet Archive's Wayback Machine to longitudinally compare configurations on 18K health websites against a top-10K control group from 2017 to 2024, finding that default-enabled tracking features such as automatic events and first-party cookies reached adoption rates up to 98.4%, that health websites show tracking of potentially sensitive information tied to booking medical appointments and button clicks associated with specific conditions such as erectile dysfunction, and that restriction features like Core Setup were configured on 34.3% of health websites versus 8.
The study introduces PixelConfig, a differential-analysis framework that reverse-engineers Meta Pixel configurations through code-patching replays and developer-account controlled experiments, and uses Internet Archive's Wayback Machine to longitudinally compare configurations on 18K health websites against a top-10K control group from 2017 to 2024, finding that default-enabled tracking features such as automatic events and first-party cookies reached adoption rates up to 98.4%, that health websites show tracking of potentially sensitive information tied to booking medical appointments and button clicks associated with specific conditions such as erectile dysfunction, and that restriction features like Core Setup were configured on 34.3% of health websites versus 8.
bioRxiv This work presents the first investigation, to the authors' knowledge, of the detectability of computationally modified sequences: the authors generate modified 16S rRNA sequences via random substitutions that pass the SILVA database quality-control inclusion criteria, and build classifiers that distinguish them from natural 16S rRNA using conserved motifs, with the best classifier achieving over 90% sensitivity and specificity on the testing set at a 5% artificial mutation rate, and one feature, gapped k-mers built from universally conserved nucleotides, conserved across all three domains of life despite relying on exact matches to patterns found in E. coli.
This work presents the first investigation, to the authors' knowledge, of the detectability of computationally modified sequences: the authors generate modified 16S rRNA sequences via random substitutions that pass the SILVA database quality-control inclusion criteria, and build classifiers that distinguish them from natural 16S rRNA using conserved motifs, with the best classifier achieving over 90% sensitivity and specificity on the testing set at a 5% artificial mutation rate, and one feature, gapped k-mers built from universally conserved nucleotides, conserved across all three domains of life despite relying on exact matches to patterns found in E. coli.
This work presents the first investigation, to the authors' knowledge, of the detectability of computationally modified sequences: the authors generate modified 16S rRNA sequences via random substitutions that pass the SILVA database quality-control inclusion criteria, and build classifiers that distinguish them from natural 16S rRNA using conserved motifs, with the best classifier achieving over 90% sensitivity and specificity on the testing set at a 5% artificial mutation rate, and one feature, gapped k-mers built from universally conserved nucleotides, conserved across all three domains of life despite relying on exact matches to patterns found in E. coli.
This work presents the first investigation, to the authors' knowledge, of the detectability of computationally modified sequences: the authors generate modified 16S rRNA sequences via random substitutions that pass the SILVA database quality-control inclusion criteria, and build classifiers that distinguish them from natural 16S rRNA using conserved motifs, with the best classifier achieving over 90% sensitivity and specificity on the testing set at a 5% artificial mutation rate, and one feature, gapped k-mers built from universally conserved nucleotides, conserved across all three domains of life despite relying on exact matches to patterns found in E. coli.
发表出处待核验 Using XMap to probe port 11434 across the full IPv4 space daily from 2025-02-14 to 2026-02-13 over 362 observed days, a Nankai University team found 152,137 cumulative exposed Ollama endpoints, with daily alive endpoints rising from 10,473 to 16,059 (+53.3%), 26.4% of IPs appearing on a single day, only 0.43%–2.90% of below-fix IPs upgrading in place across five CVE cutoffs, the top five countries/regions holding about 70% of weighted observations, and the top five ASNs all cloud or hosting providers, complemented by PTR and port-443 TLS probing of operational characteristics.
Using XMap to probe port 11434 across the full IPv4 space daily from 2025-02-14 to 2026-02-13 over 362 observed days, a Nankai University team found 152,137 cumulative exposed Ollama endpoints, with daily alive endpoints rising from 10,473 to 16,059 (+53.3%), 26.4% of IPs appearing on a single day, only 0.43%–2.90% of below-fix IPs upgrading in place across five CVE cutoffs, the top five countries/regions holding about 70% of weighted observations, and the top five ASNs all cloud or hosting providers, complemented by PTR and port-443 TLS probing of operational characteristics.
Using XMap to probe port 11434 across the full IPv4 space daily from 2025-02-14 to 2026-02-13 over 362 observed days, a Nankai University team found 152,137 cumulative exposed Ollama endpoints, with daily alive endpoints rising from 10,473 to 16,059 (+53.3%), 26.4% of IPs appearing on a single day, only 0.43%–2.90% of below-fix IPs upgrading in place across five CVE cutoffs, the top five countries/regions holding about 70% of weighted observations, and the top five ASNs all cloud or hosting providers, complemented by PTR and port-443 TLS probing of operational characteristics.
Using XMap to probe port 11434 across the full IPv4 space daily from 2025-02-14 to 2026-02-13 over 362 observed days, a Nankai University team found 152,137 cumulative exposed Ollama endpoints, with daily alive endpoints rising from 10,473 to 16,059 (+53.3%), 26.4% of IPs appearing on a single day, only 0.43%–2.90% of below-fix IPs upgrading in place across five CVE cutoffs, the top five countries/regions holding about 70% of weighted observations, and the top five ASNs all cloud or hosting providers, complemented by PTR and port-443 TLS probing of operational characteristics.
发表出处待核验 The authors present and release the IPv6 Punching Bag, a single-machine, low-memory local environment that answers ICMPv6 echo requests with per-prefix configurable response rates and address types, and use it to evaluate six dynamic target generation algorithms (6Hit, 6Sense, 6Scan, 6Tree, AddrMiner-S, and DET), finding that they only partially adhere to scanning budgets while generally respecting rate limits, that all but 6Sense fail to recognize aliased prefixes, and that 6Scan does not adapt to differing response behavior within prefixes.
The authors present and release the IPv6 Punching Bag, a single-machine, low-memory local environment that answers ICMPv6 echo requests with per-prefix configurable response rates and address types, and use it to evaluate six dynamic target generation algorithms (6Hit, 6Sense, 6Scan, 6Tree, AddrMiner-S, and DET), finding that they only partially adhere to scanning budgets while generally respecting rate limits, that all but 6Sense fail to recognize aliased prefixes, and that 6Scan does not adapt to differing response behavior within prefixes.
The authors present and release the IPv6 Punching Bag, a single-machine, low-memory local environment that answers ICMPv6 echo requests with per-prefix configurable response rates and address types, and use it to evaluate six dynamic target generation algorithms (6Hit, 6Sense, 6Scan, 6Tree, AddrMiner-S, and DET), finding that they only partially adhere to scanning budgets while generally respecting rate limits, that all but 6Sense fail to recognize aliased prefixes, and that 6Scan does not adapt to differing response behavior within prefixes.
The authors present and release the IPv6 Punching Bag, a single-machine, low-memory local environment that answers ICMPv6 echo requests with per-prefix configurable response rates and address types, and use it to evaluate six dynamic target generation algorithms (6Hit, 6Sense, 6Scan, 6Tree, AddrMiner-S, and DET), finding that they only partially adhere to scanning budgets while generally respecting rate limits, that all but 6Sense fail to recognize aliased prefixes, and that 6Scan does not adapt to differing response behavior within prefixes.
Communicable diseases intelligence (2018) Using weekly Google Trends search volumes for 2018 and 2019 compared against weekly influenza notifications from Australia's National Notifiable Disease Surveillance System (NNDSS), the study fitted four supervised regression models (elastic net, support vector regression, random forest, and feedforward neural network) independently for each state and territory except the Australian Capital Territory, for nowcast and one- and two-week-ahead predictions, finding that search volumes correlate with reported influenza rates over time, that random forest and elastic net generally performed better than the other models, that every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notificatio
Using weekly Google Trends search volumes for 2018 and 2019 compared against weekly influenza notifications from Australia's National Notifiable Disease Surveillance System (NNDSS), the study fitted four supervised regression models (elastic net, support vector regression, random forest, and feedforward neural network) independently for each state and territory except the Australian Capital Territory, for nowcast and one- and two-week-ahead predictions, finding that search volumes correlate with reported influenza rates over time, that random forest and elastic net generally performed better than the other models, that every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notificatio
Using weekly Google Trends search volumes for 2018 and 2019 compared against weekly influenza notifications from Australia's National Notifiable Disease Surveillance System (NNDSS), the study fitted four supervised regression models (elastic net, support vector regression, random forest, and feedforward neural network) independently for each state and territory except the Australian Capital Territory, for nowcast and one- and two-week-ahead predictions, finding that search volumes correlate with reported influenza rates over time, that random forest and elastic net generally performed better than the other models, that every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notificatio
Using weekly Google Trends search volumes for 2018 and 2019 compared against weekly influenza notifications from Australia's National Notifiable Disease Surveillance System (NNDSS), the study fitted four supervised regression models (elastic net, support vector regression, random forest, and feedforward neural network) independently for each state and territory except the Australian Capital Territory, for nowcast and one- and two-week-ahead predictions, finding that search volumes correlate with reported influenza rates over time, that random forest and elastic net generally performed better than the other models, that every modelled jurisdiction except the Northern Territory and Tasmania had at least two search queries with moderate to strong Pearson correlation with influenza notificatio
Journal of Machine Learning The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
发表出处待核验 Using a Puppeteer crawler on AWS Lambda, the authors ran a 40-day longitudinal measurement (March 13–April 21, 2026) over 55,393 trending Google queries across 19 topical categories and found that AI Overviews activate on 13.7% of queries overall (64.7% for question-form queries), that cited domains are more credible than co-displayed first-page results yet 29.8% do not appear on the first page at all, that 11.0% of 98,020 atomic claims are unsupported by the cited pages with omission as the dominant failure mode, and that at least 50.6% of cited pages carry display advertising.
Using a Puppeteer crawler on AWS Lambda, the authors ran a 40-day longitudinal measurement (March 13–April 21, 2026) over 55,393 trending Google queries across 19 topical categories and found that AI Overviews activate on 13.7% of queries overall (64.7% for question-form queries), that cited domains are more credible than co-displayed first-page results yet 29.8% do not appear on the first page at all, that 11.0% of 98,020 atomic claims are unsupported by the cited pages with omission as the dominant failure mode, and that at least 50.6% of cited pages carry display advertising.
Using a Puppeteer crawler on AWS Lambda, the authors ran a 40-day longitudinal measurement (March 13–April 21, 2026) over 55,393 trending Google queries across 19 topical categories and found that AI Overviews activate on 13.7% of queries overall (64.7% for question-form queries), that cited domains are more credible than co-displayed first-page results yet 29.8% do not appear on the first page at all, that 11.0% of 98,020 atomic claims are unsupported by the cited pages with omission as the dominant failure mode, and that at least 50.6% of cited pages carry display advertising.
Using a Puppeteer crawler on AWS Lambda, the authors ran a 40-day longitudinal measurement (March 13–April 21, 2026) over 55,393 trending Google queries across 19 topical categories and found that AI Overviews activate on 13.7% of queries overall (64.7% for question-form queries), that cited domains are more credible than co-displayed first-page results yet 29.8% do not appear on the first page at all, that 11.0% of 98,020 atomic claims are unsupported by the cited pages with omission as the dominant failure mode, and that at least 50.6% of cited pages carry display advertising.
bioRxiv The authors present DECIPHER, an end-to-end representation-learning framework for cell-type deconvolution that learns a domain-constant representation (Zc) for deconvolution and a domain-specific representation (Zs) to model domain-associated variation, estimates cell-type proportions from Zc via differentiable non-negative least-squares optimization, shows robust and competitive deconvolution performance across simulated datasets, experimentally generated bulk-cell mixtures, real-world datasets and multiple molecular modalities, and further shows that the learned Zc supports chronological age prediction across independent cohorts and prognostic stratification in lung adenocarcinoma.
The authors present DECIPHER, an end-to-end representation-learning framework for cell-type deconvolution that learns a domain-constant representation (Zc) for deconvolution and a domain-specific representation (Zs) to model domain-associated variation, estimates cell-type proportions from Zc via differentiable non-negative least-squares optimization, shows robust and competitive deconvolution performance across simulated datasets, experimentally generated bulk-cell mixtures, real-world datasets and multiple molecular modalities, and further shows that the learned Zc supports chronological age prediction across independent cohorts and prognostic stratification in lung adenocarcinoma.
The authors present DECIPHER, an end-to-end representation-learning framework for cell-type deconvolution that learns a domain-constant representation (Zc) for deconvolution and a domain-specific representation (Zs) to model domain-associated variation, estimates cell-type proportions from Zc via differentiable non-negative least-squares optimization, shows robust and competitive deconvolution performance across simulated datasets, experimentally generated bulk-cell mixtures, real-world datasets and multiple molecular modalities, and further shows that the learned Zc supports chronological age prediction across independent cohorts and prognostic stratification in lung adenocarcinoma.
The authors present DECIPHER, an end-to-end representation-learning framework for cell-type deconvolution that learns a domain-constant representation (Zc) for deconvolution and a domain-specific representation (Zs) to model domain-associated variation, estimates cell-type proportions from Zc via differentiable non-negative least-squares optimization, shows robust and competitive deconvolution performance across simulated datasets, experimentally generated bulk-cell mixtures, real-world datasets and multiple molecular modalities, and further shows that the learned Zc supports chronological age prediction across independent cohorts and prognostic stratification in lung adenocarcinoma.
Journal of the Physical Society of Japan Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
发表出处待核验 The study introduces a privacy measurement suite for synthetic network traffic spanning membership inference, data extraction, and network-specific attacks on identifiers, attributes, and topology, and evaluates four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) across five datasets, finding that even with black-box-only attackers and minimal training, membership inference reaches up to 0.88 TPR at FPR≤0.01 and up to 100% of network identifiers can be recovered, while anonymization and DP noise each cover only part of the risk at a utility cost.
The study introduces a privacy measurement suite for synthetic network traffic spanning membership inference, data extraction, and network-specific attacks on identifiers, attributes, and topology, and evaluates four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) across five datasets, finding that even with black-box-only attackers and minimal training, membership inference reaches up to 0.88 TPR at FPR≤0.01 and up to 100% of network identifiers can be recovered, while anonymization and DP noise each cover only part of the risk at a utility cost.
The study introduces a privacy measurement suite for synthetic network traffic spanning membership inference, data extraction, and network-specific attacks on identifiers, attributes, and topology, and evaluates four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) across five datasets, finding that even with black-box-only attackers and minimal training, membership inference reaches up to 0.88 TPR at FPR≤0.01 and up to 100% of network identifiers can be recovered, while anonymization and DP noise each cover only part of the risk at a utility cost.
The study introduces a privacy measurement suite for synthetic network traffic spanning membership inference, data extraction, and network-specific attacks on identifiers, attributes, and topology, and evaluates four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) across five datasets, finding that even with black-box-only attackers and minimal training, membership inference reaches up to 0.88 TPR at FPR≤0.01 and up to 100% of network identifiers can be recovered, while anonymization and DP noise each cover only part of the risk at a utility cost.
bioRxiv The authors developed SpatialTRACE, comprising graph- and image-based models: SpatialTRACE-Graph combines gene-expression profiles with spatial neighborhoods to predict crypt-villus and epithelial-distance axis coordinates and to identify Peyer's patches in mouse small-intestine sections using as few as 10 annotated training villi, while SpatialTRACE-Image, a multiscale vision transformer that learns from the coordinate and region predictions generated by SpatialTRACE-Graph, predicts the same anatomical axis coordinates and regions across entire tissue images from DAPI alone and was applied to immunofluorescence images to map antigen-specific P14 CD8 T cells responding to acute systemic LCMV Armstrong infection in the small intestine, showing that a retinoic acid receptor inhibitor-treated
The authors developed SpatialTRACE, comprising graph- and image-based models: SpatialTRACE-Graph combines gene-expression profiles with spatial neighborhoods to predict crypt-villus and epithelial-distance axis coordinates and to identify Peyer's patches in mouse small-intestine sections using as few as 10 annotated training villi, while SpatialTRACE-Image, a multiscale vision transformer that learns from the coordinate and region predictions generated by SpatialTRACE-Graph, predicts the same anatomical axis coordinates and regions across entire tissue images from DAPI alone and was applied to immunofluorescence images to map antigen-specific P14 CD8 T cells responding to acute systemic LCMV Armstrong infection in the small intestine, showing that a retinoic acid receptor inhibitor-treated
The authors developed SpatialTRACE, comprising graph- and image-based models: SpatialTRACE-Graph combines gene-expression profiles with spatial neighborhoods to predict crypt-villus and epithelial-distance axis coordinates and to identify Peyer's patches in mouse small-intestine sections using as few as 10 annotated training villi, while SpatialTRACE-Image, a multiscale vision transformer that learns from the coordinate and region predictions generated by SpatialTRACE-Graph, predicts the same anatomical axis coordinates and regions across entire tissue images from DAPI alone and was applied to immunofluorescence images to map antigen-specific P14 CD8 T cells responding to acute systemic LCMV Armstrong infection in the small intestine, showing that a retinoic acid receptor inhibitor-treated
The authors developed SpatialTRACE, comprising graph- and image-based models: SpatialTRACE-Graph combines gene-expression profiles with spatial neighborhoods to predict crypt-villus and epithelial-distance axis coordinates and to identify Peyer's patches in mouse small-intestine sections using as few as 10 annotated training villi, while SpatialTRACE-Image, a multiscale vision transformer that learns from the coordinate and region predictions generated by SpatialTRACE-Graph, predicts the same anatomical axis coordinates and regions across entire tissue images from DAPI alone and was applied to immunofluorescence images to map antigen-specific P14 CD8 T cells responding to acute systemic LCMV Armstrong infection in the small intestine, showing that a retinoic acid receptor inhibitor-treated
bioRxiv The study developed a tool that detects and compares insertion sequence insertions from short reads without a reference genome, applied it to 10,000 strains of the Mycobacterium tuberculosis complex (MTBC), and combined it with ancestral state reconstruction on presence-absence patterns to describe the distribution of IS6110 copy numbers (from 1 in some clades to more than 30 in strains of La3 (M.
The study developed a tool that detects and compares insertion sequence insertions from short reads without a reference genome, applied it to 10,000 strains of the Mycobacterium tuberculosis complex (MTBC), and combined it with ancestral state reconstruction on presence-absence patterns to describe the distribution of IS6110 copy numbers (from 1 in some clades to more than 30 in strains of La3 (M.
The study developed a tool that detects and compares insertion sequence insertions from short reads without a reference genome, applied it to 10,000 strains of the Mycobacterium tuberculosis complex (MTBC), and combined it with ancestral state reconstruction on presence-absence patterns to describe the distribution of IS6110 copy numbers (from 1 in some clades to more than 30 in strains of La3 (M.
The study developed a tool that detects and compares insertion sequence insertions from short reads without a reference genome, applied it to 10,000 strains of the Mycobacterium tuberculosis complex (MTBC), and combined it with ancestral state reconstruction on presence-absence patterns to describe the distribution of IS6110 copy numbers (from 1 in some clades to more than 30 in strains of La3 (M.
发表出处待核验 The authors introduce GateScope, a lightweight black-box auditing framework that uses only public APIs to evaluate LLM API gateways along response content, multi-turn conversation consistency, billing accuracy, and latency characteristics; controlled validation on official endpoints yields an average F1 of 0.968±0.085 across 24 models, and auditing 10 commercial gateways reveals identification rates as low as 13.09% for gpt-5, degraded multi-turn memory checkpoints, a 62.8% billing gap for o*ey on gpt-4o, and markedly higher latency variation for b*ie.
The authors introduce GateScope, a lightweight black-box auditing framework that uses only public APIs to evaluate LLM API gateways along response content, multi-turn conversation consistency, billing accuracy, and latency characteristics; controlled validation on official endpoints yields an average F1 of 0.968±0.085 across 24 models, and auditing 10 commercial gateways reveals identification rates as low as 13.09% for gpt-5, degraded multi-turn memory checkpoints, a 62.8% billing gap for o*ey on gpt-4o, and markedly higher latency variation for b*ie.
The authors introduce GateScope, a lightweight black-box auditing framework that uses only public APIs to evaluate LLM API gateways along response content, multi-turn conversation consistency, billing accuracy, and latency characteristics; controlled validation on official endpoints yields an average F1 of 0.968±0.085 across 24 models, and auditing 10 commercial gateways reveals identification rates as low as 13.09% for gpt-5, degraded multi-turn memory checkpoints, a 62.8% billing gap for o*ey on gpt-4o, and markedly higher latency variation for b*ie.
The authors introduce GateScope, a lightweight black-box auditing framework that uses only public APIs to evaluate LLM API gateways along response content, multi-turn conversation consistency, billing accuracy, and latency characteristics; controlled validation on official endpoints yields an average F1 of 0.968±0.085 across 24 models, and auditing 10 commercial gateways reveals identification rates as low as 13.09% for gpt-5, degraded multi-turn memory checkpoints, a 62.8% billing gap for o*ey on gpt-4o, and markedly higher latency variation for b*ie.
Claude 产品博客 Asana Chief Product Officer Arnab Bose describes how Asana runs AI agents as "AI teammates" inside its existing Work Graph: agents are built around roles (such as content writer, insights analyst, project manager, work intake specialist, campaign analyst, or campaign coordinator) with pre-built skills and integrations, distinct identities and controlled permissions, effective access bounded by the permissions of the person who triggers them, and coachable shared memory in which only admins and editors can commit feedback to permanent memory; the piece gives three deployments—Slack channel questions turned into Asana tasks and routed to product backlog or enablement documentation updates, an At-Risk Renewal agent producing a daily global renewal-risk digest for executives and regional leade
Asana Chief Product Officer Arnab Bose describes how Asana runs AI agents as "AI teammates" inside its existing Work Graph: agents are built around roles (such as content writer, insights analyst, project manager, work intake specialist, campaign analyst, or campaign coordinator) with pre-built skills and integrations, distinct identities and controlled permissions, effective access bounded by the permissions of the person who triggers them, and coachable shared memory in which only admins and editors can commit feedback to permanent memory; the piece gives three deployments—Slack channel questions turned into Asana tasks and routed to product backlog or enablement documentation updates, an At-Risk Renewal agent producing a daily global renewal-risk digest for executives and regional leade
Asana Chief Product Officer Arnab Bose describes how Asana runs AI agents as "AI teammates" inside its existing Work Graph: agents are built around roles (such as content writer, insights analyst, project manager, work intake specialist, campaign analyst, or campaign coordinator) with pre-built skills and integrations, distinct identities and controlled permissions, effective access bounded by the permissions of the person who triggers them, and coachable shared memory in which only admins and editors can commit feedback to permanent memory; the piece gives three deployments—Slack channel questions turned into Asana tasks and routed to product backlog or enablement documentation updates, an At-Risk Renewal agent producing a daily global renewal-risk digest for executives and regional leade
Asana Chief Product Officer Arnab Bose describes how Asana runs AI agents as "AI teammates" inside its existing Work Graph: agents are built around roles (such as content writer, insights analyst, project manager, work intake specialist, campaign analyst, or campaign coordinator) with pre-built skills and integrations, distinct identities and controlled permissions, effective access bounded by the permissions of the person who triggers them, and coachable shared memory in which only admins and editors can commit feedback to permanent memory; the piece gives three deployments—Slack channel questions turned into Asana tasks and routed to product backlog or enablement documentation updates, an At-Risk Renewal agent producing a daily global renewal-risk digest for executives and regional leade
bioRxiv On a four-way mood and psychosis classification task over 1,520 subjects from three studies and 14 acquisition sites, the work shows that marginal split-conformal calibration reached 0.9000 empirical coverage against a nominal 0.90 while healthy controls were covered at 0.941 and schizoaffective disorder at 0.818, and that class-conditional (Mondrian) calibration reduced this 12.3-point disparity to 0.4 points at a cost of 0.07 labels in mean set size (under 3%), making set size at matched coverage interpretable as a property of the subject and separating subjects into confident, boundary, ambiguous and unresolved strata, with the proportion independently flagged as label-ambiguous by a structural-MRI model rising monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
On a four-way mood and psychosis classification task over 1,520 subjects from three studies and 14 acquisition sites, the work shows that marginal split-conformal calibration reached 0.9000 empirical coverage against a nominal 0.90 while healthy controls were covered at 0.941 and schizoaffective disorder at 0.818, and that class-conditional (Mondrian) calibration reduced this 12.3-point disparity to 0.4 points at a cost of 0.07 labels in mean set size (under 3%), making set size at matched coverage interpretable as a property of the subject and separating subjects into confident, boundary, ambiguous and unresolved strata, with the proportion independently flagged as label-ambiguous by a structural-MRI model rising monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
On a four-way mood and psychosis classification task over 1,520 subjects from three studies and 14 acquisition sites, the work shows that marginal split-conformal calibration reached 0.9000 empirical coverage against a nominal 0.90 while healthy controls were covered at 0.941 and schizoaffective disorder at 0.818, and that class-conditional (Mondrian) calibration reduced this 12.3-point disparity to 0.4 points at a cost of 0.07 labels in mean set size (under 3%), making set size at matched coverage interpretable as a property of the subject and separating subjects into confident, boundary, ambiguous and unresolved strata, with the proportion independently flagged as label-ambiguous by a structural-MRI model rising monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
On a four-way mood and psychosis classification task over 1,520 subjects from three studies and 14 acquisition sites, the work shows that marginal split-conformal calibration reached 0.9000 empirical coverage against a nominal 0.90 while healthy controls were covered at 0.941 and schizoaffective disorder at 0.818, and that class-conditional (Mondrian) calibration reduced this 12.3-point disparity to 0.4 points at a cost of 0.07 labels in mean set size (under 3%), making set size at matched coverage interpretable as a property of the subject and separating subjects into confident, boundary, ambiguous and unresolved strata, with the proportion independently flagged as label-ambiguous by a structural-MRI model rising monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
bioRxiv The authors present TRIDENT-2, a multimodal artificial intelligence model trained on 560,780 toxicity assays spanning 82,775 chemicals, 6,793 species, and multiple exposure scenarios to predict chemical toxicity across evolutionarily diverse eukaryotic species, reporting an average median absolute error of 1.76 to 3.80 and remaining accurate across broad chemical and taxonomic distances, which allows toxicity assessment for species and chemicals beyond current experimental evidence.
The authors present TRIDENT-2, a multimodal artificial intelligence model trained on 560,780 toxicity assays spanning 82,775 chemicals, 6,793 species, and multiple exposure scenarios to predict chemical toxicity across evolutionarily diverse eukaryotic species, reporting an average median absolute error of 1.76 to 3.80 and remaining accurate across broad chemical and taxonomic distances, which allows toxicity assessment for species and chemicals beyond current experimental evidence.
The authors present TRIDENT-2, a multimodal artificial intelligence model trained on 560,780 toxicity assays spanning 82,775 chemicals, 6,793 species, and multiple exposure scenarios to predict chemical toxicity across evolutionarily diverse eukaryotic species, reporting an average median absolute error of 1.76 to 3.80 and remaining accurate across broad chemical and taxonomic distances, which allows toxicity assessment for species and chemicals beyond current experimental evidence.
The authors present TRIDENT-2, a multimodal artificial intelligence model trained on 560,780 toxicity assays spanning 82,775 chemicals, 6,793 species, and multiple exposure scenarios to predict chemical toxicity across evolutionarily diverse eukaryotic species, reporting an average median absolute error of 1.76 to 3.80 and remaining accurate across broad chemical and taxonomic distances, which allows toxicity assessment for species and chemicals beyond current experimental evidence.
bioRxiv The authors present and open-source Murmurent, shared software that sits beneath agentic AI for biomedical labs, offering multi-member project and "choreography" infrastructure, specialized agents for typical biomedical data science tasks, tiered memory, traceability records, SOP and data-governance enforcement, and multi-user collaboration, and they use the system to identify putative inhibitors of Peptidyl-prolyl cis-trans Isomerase NIMA-interacting 1 (Pin1), describing several approaches and the results they yield.
The authors present and open-source Murmurent, shared software that sits beneath agentic AI for biomedical labs, offering multi-member project and "choreography" infrastructure, specialized agents for typical biomedical data science tasks, tiered memory, traceability records, SOP and data-governance enforcement, and multi-user collaboration, and they use the system to identify putative inhibitors of Peptidyl-prolyl cis-trans Isomerase NIMA-interacting 1 (Pin1), describing several approaches and the results they yield.
The authors present and open-source Murmurent, shared software that sits beneath agentic AI for biomedical labs, offering multi-member project and "choreography" infrastructure, specialized agents for typical biomedical data science tasks, tiered memory, traceability records, SOP and data-governance enforcement, and multi-user collaboration, and they use the system to identify putative inhibitors of Peptidyl-prolyl cis-trans Isomerase NIMA-interacting 1 (Pin1), describing several approaches and the results they yield.
The authors present and open-source Murmurent, shared software that sits beneath agentic AI for biomedical labs, offering multi-member project and "choreography" infrastructure, specialized agents for typical biomedical data science tasks, tiered memory, traceability records, SOP and data-governance enforcement, and multi-user collaboration, and they use the system to identify putative inhibitors of Peptidyl-prolyl cis-trans Isomerase NIMA-interacting 1 (Pin1), describing several approaches and the results they yield.
发表出处待核验 The authors built an automated detection system combining lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification, applied it to 6,094,475 public URLs collected from VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt paste sites, and identified 12,331 potential exposures across authentication, financial, personal, and document-related categories, including 26 live password-reset links, 83 API keys, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes, and 62 JWTs lacking expiration constraints.
The authors built an automated detection system combining lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification, applied it to 6,094,475 public URLs collected from VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt paste sites, and identified 12,331 potential exposures across authentication, financial, personal, and document-related categories, including 26 live password-reset links, 83 API keys, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes, and 62 JWTs lacking expiration constraints.
The authors built an automated detection system combining lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification, applied it to 6,094,475 public URLs collected from VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt paste sites, and identified 12,331 potential exposures across authentication, financial, personal, and document-related categories, including 26 live password-reset links, 83 API keys, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes, and 62 JWTs lacking expiration constraints.
The authors built an automated detection system combining lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification, applied it to 6,094,475 public URLs collected from VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt paste sites, and identified 12,331 potential exposures across authentication, financial, personal, and document-related categories, including 26 live password-reset links, 83 API keys, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes, and 62 JWTs lacking expiration constraints.
CSIAM Transactions on Applied Mathematics The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
The authors propose a Derivative-informed Graph Convolutional Autoencoder (DiGCA) phase classifier that feeds both the Lifshitz-Petrich model solutions and their derivatives (the nonlocal term G(φ)) into a graph convolutional autoencoder for dimensionality reduction, then classifies with a fully connected neural network, generating phase diagrams over the parameter domain [−0.01,0.05]×[0,1] with over 98% classification accuracy, roughly two orders of magnitude faster than MCMS-RBM, and remaining stable under up to 10% additive white noise.
bioRxiv The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
MIT Technology Review An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
An MIT Technology Review investigation documented over a thousand people who moved through areas watched by surveillance towers along the southern US border without being reached or apprehended and who ultimately died there, some under newly installed AI-powered towers designed to spot people automatically, revealing repeated failures of the virtual wall's basic security promise.
Microsoft Research Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Microsoft Research Asia – Singapore reviews its first year since opening on July 24, 2025: it pursues a "research-to-impact flywheel" across four pillars (next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development), launched nine new projects with the National University of Singapore and Nanyang Technological University, signed a five-year Framework Research Agreement with NUS, reached more than 300 students through its summer school since 2025, and saw its first Industrial Postgraduate Programme student Qiming Huang have a RobotSeg paper selected for an oral presentation at CVPR 2026, with the second year focused on scaling what works.
Terence Tao blog RSS A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
A coalition of Caltech mathematicians—including the original Mathathon organizers, two coauthors of the Open Letter about the Mathathon, and other community members—issued a joint statement announcing that Mathathon is being redesigned around the theme "Old Problems, New Proofs": participants pick a solved problem with an unintuitive solution, learn as much as possible in 40 hours and present findings to peers, then take two months to develop an alternative proof or exposition, submitting an explainer in any form plus a GitHub repository holding whatever produced it (LLM chat histories, visualization code, and more); the event runs November 13–15, is co-organized with the Foundation for Science and AI Research (SAIR), has XTX Markets as lead donor, and no longer takes sponsorships from dev
IEEE Spectrum IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
IEEE TryEngineering, through the Lerner Publishing Group, has introduced the six-book series "Tomorrow's Technology With TryEngineering, Powered by IEEE" for children ages 8 to 12, with each book covering artificial intelligence, communication technology, electric vehicles, ocean engineering, semiconductors, or signal processing, combining age-appropriate explanations, real-world examples, and design challenges, based on ebooks and videos available at tryengineering.org and developed with IEEE Communications, Computer, and Oceanic Engineering societies and the Transportation Electrification Council.
MIT Technology Review Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
Anthropic announced that 950 Claude agents in its molecular biology lab flagged, within 21 hours, a repeating pattern surrounding a known enzyme in a large library of DNA sequences, saying the pattern had not been catalogued before, while some biologists argue that finding gene clusters and repeats is often the easy part and that the real discovery lies in figuring out what the system actually does, and a University of Copenhagen researcher says his team had already found the pattern.
arXiv The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
arXiv The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
arXiv Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
arXiv The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
arXiv The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
The work introduces an auditable protocol for homogeneous three-agent LLM debate that decomposes final accuracy into preserved, collapse, correction, and unrepaired transitions with collapse onset and signed intervention utility, identifies 253 collapses across 6,925 MMLU-Pro debates, and shows via replay that a leave-one-model-out probe-gated freeze prevents 29 collapses while losing 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy.
arXiv EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
arXiv The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
The work introduces VisionHOPE, which formulates a visual backbone as a self-modifying learning system in which five coupled memories for content, key, value, learning rate, and retention co-evolve along each scan, and pairs this with a stability-matched step-size control (a soft cap on self-referential injection plus a spectral clamp on the retained memory transition) that is proved to yield non-expansive memory dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.
arXiv The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
arXiv TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.
arXiv The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
arXiv EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
EvolvingAvatar introduces a causal interactive 3D head generator that adapts at inference time via test-time training on user face video and dyadic audio, using a Dyadic Context Prediction self-supervised objective with persistent fast weights and transient jaw adaptation to produce speaking and listening motion, and releases InterHead-Bench, a 455.95-hour benchmark; experiments show improved conversational motion statistics over strong baselines and improvement as conversations unfold on the hardest out-of-distribution split, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
arXiv NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
NVAlign is a direct-gradient post-training framework for non-verbal control in continuous autoregressive flow-matching TTS: it first supervised-fine-tunes TTS models and a Qwen3-Omni-30B-based NV-ASR model on NVV-annotated speech, then freezes the NV-ASR as a reward model and uses a two-step gradient surrogate to backpropagate target-tag probabilities through the flow-matching sampler, jointly updating the autoregressive backbone and acoustic flow head under fidelity penalties and reference-velocity regularization; on NVV-SuperBench and blinded listening evaluations it improves tag-following accuracy over SFT and Flow-GRPO baselines on VoxCPM2, dots.tts, and an English production system.
arXiv Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
Nature News Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
Nature reports that Vita, a new English-language open-access journal jointly funded by the Higher Education Press and Westlake University, publishes biological-sciences papers, released its first print issue in July and charges no article-processing fees, while China's Journal Excellence Action Plan tiers 450 journals for up to 1.5 million yuan each per year over five years and supports 13 academic publishers, aiming to build globally influential Chinese-owned journals.
arXiv Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
arXiv Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
arXiv The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
arXiv The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
arXiv The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
The work introduces the Adaptive Consistency Graph (ACG), which incrementally organizes execution evidence and its provenance into a persistent graph and builds a temporary requirement-centered view for each decision under a bounded context budget, raising GPT-5.6-luna's equal-weight average success across three benchmarks from 44.5% with ReAct to 50.2% and BrowseComp-Plus from 62.4% to 73.5%, without replacing the base planner or tool executor.
arXiv The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
The work casts quadrilateral block decomposition as a Markov decision process over a half-edge mesh, uses the vertex-irregularity lower bound implied by the discrete Gauss–Bonnet identity (called par) as both reward target and termination test, and trains via behaviour cloning on trivially constructible optimal meshes followed by PPO, producing an all-quadrilateral mesh on all 96 held-out domains, a usable one on 95.7 on average and a provably optimal one on 90, whereas Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none.
arXiv ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
arXiv The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
The work introduces a forward-pass-matched diagnostic protocol for reward-guided reasoning in discrete diffusion language models and shows, on Dream-7B and LLaDA-8B-Base, that deterministic top-k process-reward-model guidance trails independent sampling plus task-matched outcome-reward-model reranking on GSM8K, MATH, and MBPP, decomposing the gap into candidate-pool damage from guidance and terminal selection quality.
arXiv The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
arXiv WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
arXiv The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
The work introduces AnswerMap: the image is cut into K row and K column bands, each shown alone to a frozen VLM with a yes/no question about the query, and the outer product of the row and column "yes" posteriors gives a query-conditioned spatial map on which a fixed read-out (e.g., expectation, maximum) derives continuous outputs natively; across four models and three query distributions the map agrees with the model's own generated point at AUC 0.85 versus 0.38 for attention, deleting its region flips 53% of correct answers versus 19% for attention, and its maximum flags hallucinated objects, its expectation localizes when the model's own pointing fails, and its top-mass crop fixes about half of the model's wrong answers.
arXiv The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
arXiv The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.
arXiv Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
arXiv This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
This work studies the test-time scaling of looped transformers through post-training, finds that existing looped models have steeper slopes yet underperform the non-looped baseline at matched compute, and proposes TaH2, which jointly post-trains the backbone and an iteration decider with lookahead depth supervision to label online which tokens benefit from further iterations, improving the AIME accuracy-compute slope from 1.79 to 2.74 (53%) and exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute.
arXiv The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
The work presents KVCMAS, an online KV cache correction framework for prompt-specialized multi-agent systems that represents cross-agent cache deviations as compact low-rank states and chains corrections along the agent workflow without an additional context-free reference prefill, matching or improving the accuracy of prior KV cache sharing methods across language and vision-language workloads while achieving a 2.0 TTFT speedup over inference without KV cache sharing at 32K shared tokens and 8 QPS and reducing peak GPU memory by up to 3.7 relative to the prior correction method KVComm.
arXiv The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
The work first audits existing latent communication interfaces: across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points even when communication adds 15.44 points over the receiver alone, indicating the gain can come from interface adaptation rather than message content; it then introduces Draft-KV, which sends the key-value states formed while the sharer drafts an answer through linear projections into a side memory read by a gated attention branch, trained progressively from message reconstruction to answer supervision under a one-sided guard on harm from mismatched messages, with both models frozen and only 1.05M interface parameters trained (348x fewer than C2C); with a Qwen3-8B sharer and a frozen Qwen2.5-0.
arXiv Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
arXiv Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
arXiv Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
arXiv The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
arXiv The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
arXiv YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
arXiv The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
arXiv FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
arXiv KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
arXiv The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
arXiv The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
arXiv GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
arXiv The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
arXiv CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
arXiv The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
arXiv The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
arXiv The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
arXiv The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
arXiv The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
arXiv Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
arXiv The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
arXiv The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
arXiv The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
arXiv The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
arXiv The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
arXiv Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
arXiv InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
arXiv REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
arXiv The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
arXiv SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
arXiv The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
arXiv VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
arXiv QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
Google AI 与 Gemini 产品博客 Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
arXiv The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
arXiv The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
arXiv This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
arXiv The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
arXiv Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
arXiv The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
arXiv The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
arXiv Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
arXiv The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
arXiv The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
arXiv The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
arXiv The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
arXiv The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
arXiv The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
arXiv Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
arXiv The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
arXiv RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
arXiv The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
arXiv RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
arXiv Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
arXiv This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
arXiv TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
arXiv The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
arXiv In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
arXiv The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
Harvard University Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
arXiv The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
arXiv The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
arXiv SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
arXiv The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
arXiv The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
arXiv The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
Science News Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
arXiv The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
arXiv The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
arXiv The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
arXiv Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
arXiv The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
arXiv VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
arXiv The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.