Journal of Applied Health Sciences and Medicine This case report describes a 23-year-old man who presented with a 3×4 cm cystic left-sided neck swelling that moved with deglutition but not clearly with tongue protrusion; ultrasound and contrast-enhanced neck CT showed a cystic lesion below the hyoid and above the thyroid cartilage extending laterally to the left, FNAC suggested a benign cystic lesion possibly a thyroglossal duct cyst, and the patient underwent a Sistrunk procedure removing the cyst, tract and body of the hyoid, with an uneventful postoperative course and histopathology confirming a left thyroglossal duct cyst.
This case report describes a 23-year-old man who presented with a 3×4 cm cystic left-sided neck swelling that moved with deglutition but not clearly with tongue protrusion; ultrasound and contrast-enhanced neck CT showed a cystic lesion below the hyoid and above the thyroid cartilage extending laterally to the left, FNAC suggested a benign cystic lesion possibly a thyroglossal duct cyst, and the patient underwent a Sistrunk procedure removing the cyst, tract and body of the hyoid, with an uneventful postoperative course and histopathology confirming a left thyroglossal duct cyst.
This case report describes a 23-year-old man who presented with a 3×4 cm cystic left-sided neck swelling that moved with deglutition but not clearly with tongue protrusion; ultrasound and contrast-enhanced neck CT showed a cystic lesion below the hyoid and above the thyroid cartilage extending laterally to the left, FNAC suggested a benign cystic lesion possibly a thyroglossal duct cyst, and the patient underwent a Sistrunk procedure removing the cyst, tract and body of the hyoid, with an uneventful postoperative course and histopathology confirming a left thyroglossal duct cyst.
This case report describes a 23-year-old man who presented with a 3×4 cm cystic left-sided neck swelling that moved with deglutition but not clearly with tongue protrusion; ultrasound and contrast-enhanced neck CT showed a cystic lesion below the hyoid and above the thyroid cartilage extending laterally to the left, FNAC suggested a benign cystic lesion possibly a thyroglossal duct cyst, and the patient underwent a Sistrunk procedure removing the cyst, tract and body of the hyoid, with an uneventful postoperative course and histopathology confirming a left thyroglossal duct cyst.
Natural Sciences and Applied Technology This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
Natural Sciences and Applied Technology The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
Claude 产品博客 Carl Johnson, a sales development leader at Anthropic, describes how his team built a buying agent on Claude Managed Agents (beta), deployed on the Contact Sales and Pricing pages, inside the product, and in email, which now holds thousands of conversations a day and can take buyers through checkout, turning leads into opportunities more than twice as often as the old form, closing about five days faster, and cutting by about half the share of conversations that needed a person to close.
Carl Johnson, a sales development leader at Anthropic, describes how his team built a buying agent on Claude Managed Agents (beta), deployed on the Contact Sales and Pricing pages, inside the product, and in email, which now holds thousands of conversations a day and can take buyers through checkout, turning leads into opportunities more than twice as often as the old form, closing about five days faster, and cutting by about half the share of conversations that needed a person to close.
Carl Johnson, a sales development leader at Anthropic, describes how his team built a buying agent on Claude Managed Agents (beta), deployed on the Contact Sales and Pricing pages, inside the product, and in email, which now holds thousands of conversations a day and can take buyers through checkout, turning leads into opportunities more than twice as often as the old form, closing about five days faster, and cutting by about half the share of conversations that needed a person to close.
Carl Johnson, a sales development leader at Anthropic, describes how his team built a buying agent on Claude Managed Agents (beta), deployed on the Contact Sales and Pricing pages, inside the product, and in email, which now holds thousands of conversations a day and can take buyers through checkout, turning leads into opportunities more than twice as often as the old form, closing about five days faster, and cutting by about half the share of conversations that needed a person to close.
The FASEB Journal Combining limma differential analysis, WGCNA, a GeneCards oxidative-stress gene set, PPI networks, and three machine learning methods (LASSO, random forest, SVM-RFE), the study narrowed obstructive sleep apnea (OSA) adipose transcriptomes to 57 shared differentially expressed genes and two hub genes, JAK2 and ANXA5, then used single-cell sequencing, scTenifoldKnk virtual knockout, immune deconvolution, RT-qPCR, and Western blotting to show that JAK2 is significantly upregulated and ANXA5 significantly downregulated in OSA, that both are enriched in monocytes, and that CPAP treatment lowers JAK2 while raising ANXA5.
Combining limma differential analysis, WGCNA, a GeneCards oxidative-stress gene set, PPI networks, and three machine learning methods (LASSO, random forest, SVM-RFE), the study narrowed obstructive sleep apnea (OSA) adipose transcriptomes to 57 shared differentially expressed genes and two hub genes, JAK2 and ANXA5, then used single-cell sequencing, scTenifoldKnk virtual knockout, immune deconvolution, RT-qPCR, and Western blotting to show that JAK2 is significantly upregulated and ANXA5 significantly downregulated in OSA, that both are enriched in monocytes, and that CPAP treatment lowers JAK2 while raising ANXA5.
Combining limma differential analysis, WGCNA, a GeneCards oxidative-stress gene set, PPI networks, and three machine learning methods (LASSO, random forest, SVM-RFE), the study narrowed obstructive sleep apnea (OSA) adipose transcriptomes to 57 shared differentially expressed genes and two hub genes, JAK2 and ANXA5, then used single-cell sequencing, scTenifoldKnk virtual knockout, immune deconvolution, RT-qPCR, and Western blotting to show that JAK2 is significantly upregulated and ANXA5 significantly downregulated in OSA, that both are enriched in monocytes, and that CPAP treatment lowers JAK2 while raising ANXA5.
Combining limma differential analysis, WGCNA, a GeneCards oxidative-stress gene set, PPI networks, and three machine learning methods (LASSO, random forest, SVM-RFE), the study narrowed obstructive sleep apnea (OSA) adipose transcriptomes to 57 shared differentially expressed genes and two hub genes, JAK2 and ANXA5, then used single-cell sequencing, scTenifoldKnk virtual knockout, immune deconvolution, RT-qPCR, and Western blotting to show that JAK2 is significantly upregulated and ANXA5 significantly downregulated in OSA, that both are enriched in monocytes, and that CPAP treatment lowers JAK2 while raising ANXA5.
Natural Sciences and Applied Technology The work introduces an interpretable Fuzzy Rank-H2O AutoML framework in which a Fuzzy Rank Feature Selection (FRFS) algorithm picks predictors by combining statistical significance, information gain, clinical importance, and uncertainty, H2O AutoML then automatically builds ensemble models, and SHAP and LIME supply global and patient-level explanations; evaluated on two public cervical cancer datasets with stratified five-fold cross-validation where SMOTE is applied only to training folds to avoid information leakage, it reports a highest accuracy of 95.1% and AUC of 98.1%, outperforming traditional machine learning models and the baseline H2O AutoML framework.
The work introduces an interpretable Fuzzy Rank-H2O AutoML framework in which a Fuzzy Rank Feature Selection (FRFS) algorithm picks predictors by combining statistical significance, information gain, clinical importance, and uncertainty, H2O AutoML then automatically builds ensemble models, and SHAP and LIME supply global and patient-level explanations; evaluated on two public cervical cancer datasets with stratified five-fold cross-validation where SMOTE is applied only to training folds to avoid information leakage, it reports a highest accuracy of 95.1% and AUC of 98.1%, outperforming traditional machine learning models and the baseline H2O AutoML framework.
The work introduces an interpretable Fuzzy Rank-H2O AutoML framework in which a Fuzzy Rank Feature Selection (FRFS) algorithm picks predictors by combining statistical significance, information gain, clinical importance, and uncertainty, H2O AutoML then automatically builds ensemble models, and SHAP and LIME supply global and patient-level explanations; evaluated on two public cervical cancer datasets with stratified five-fold cross-validation where SMOTE is applied only to training folds to avoid information leakage, it reports a highest accuracy of 95.1% and AUC of 98.1%, outperforming traditional machine learning models and the baseline H2O AutoML framework.
The work introduces an interpretable Fuzzy Rank-H2O AutoML framework in which a Fuzzy Rank Feature Selection (FRFS) algorithm picks predictors by combining statistical significance, information gain, clinical importance, and uncertainty, H2O AutoML then automatically builds ensemble models, and SHAP and LIME supply global and patient-level explanations; evaluated on two public cervical cancer datasets with stratified five-fold cross-validation where SMOTE is applied only to training folds to avoid information leakage, it reports a highest accuracy of 95.1% and AUC of 98.1%, outperforming traditional machine learning models and the baseline H2O AutoML framework.
NVIDIA Technical Blog In an experience report, the NVIDIA team describes how it built the open source TensorRT Model Connect: a C++ collection of model-family-owned reference implementations on top of TensorRT that turn supported Hugging Face or local checkpoints into versioned .bundle artifacts and expose task-oriented native C++ APIs for text, vision, audio, diffusion, segmentation, embedding, forecasting, and other workloads; as of the public July 29, 2026 release comparison the project covered 128 model families tested on NVIDIA GB300, and the team derives an operational "AI native" practice centered on parallel decomposable work, model-family isolation, reversible changes, and GPU-backed automated validation.
In an experience report, the NVIDIA team describes how it built the open source TensorRT Model Connect: a C++ collection of model-family-owned reference implementations on top of TensorRT that turn supported Hugging Face or local checkpoints into versioned .bundle artifacts and expose task-oriented native C++ APIs for text, vision, audio, diffusion, segmentation, embedding, forecasting, and other workloads; as of the public July 29, 2026 release comparison the project covered 128 model families tested on NVIDIA GB300, and the team derives an operational "AI native" practice centered on parallel decomposable work, model-family isolation, reversible changes, and GPU-backed automated validation.
In an experience report, the NVIDIA team describes how it built the open source TensorRT Model Connect: a C++ collection of model-family-owned reference implementations on top of TensorRT that turn supported Hugging Face or local checkpoints into versioned .bundle artifacts and expose task-oriented native C++ APIs for text, vision, audio, diffusion, segmentation, embedding, forecasting, and other workloads; as of the public July 29, 2026 release comparison the project covered 128 model families tested on NVIDIA GB300, and the team derives an operational "AI native" practice centered on parallel decomposable work, model-family isolation, reversible changes, and GPU-backed automated validation.
In an experience report, the NVIDIA team describes how it built the open source TensorRT Model Connect: a C++ collection of model-family-owned reference implementations on top of TensorRT that turn supported Hugging Face or local checkpoints into versioned .bundle artifacts and expose task-oriented native C++ APIs for text, vision, audio, diffusion, segmentation, embedding, forecasting, and other workloads; as of the public July 29, 2026 release comparison the project covered 128 model families tested on NVIDIA GB300, and the team derives an operational "AI native" practice centered on parallel decomposable work, model-family isolation, reversible changes, and GPU-backed automated validation.
The latest research from Google Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
NVIDIA Technical Blog NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
Harvard Gazette In a panel discussion hosted by the Berkman Klein Center, where Alex Pascal opened by asking what we want from AI, Jeff Dunn, Amit Goldenberg, and Avijit Ghosh discussed AI's double-edged nature, a continuum from tool to full-fledged named agent with personality, and the shortcomings of current "benchmark maxing" evaluation; Ghosh proposed an evaluation system that measures both system behavior and impact on people and public participation in reporting and assessment, for example incentivizing companies with liability relief if they fix a reported problem within 60 days.
In a panel discussion hosted by the Berkman Klein Center, where Alex Pascal opened by asking what we want from AI, Jeff Dunn, Amit Goldenberg, and Avijit Ghosh discussed AI's double-edged nature, a continuum from tool to full-fledged named agent with personality, and the shortcomings of current "benchmark maxing" evaluation; Ghosh proposed an evaluation system that measures both system behavior and impact on people and public participation in reporting and assessment, for example incentivizing companies with liability relief if they fix a reported problem within 60 days.
In a panel discussion hosted by the Berkman Klein Center, where Alex Pascal opened by asking what we want from AI, Jeff Dunn, Amit Goldenberg, and Avijit Ghosh discussed AI's double-edged nature, a continuum from tool to full-fledged named agent with personality, and the shortcomings of current "benchmark maxing" evaluation; Ghosh proposed an evaluation system that measures both system behavior and impact on people and public participation in reporting and assessment, for example incentivizing companies with liability relief if they fix a reported problem within 60 days.
In a panel discussion hosted by the Berkman Klein Center, where Alex Pascal opened by asking what we want from AI, Jeff Dunn, Amit Goldenberg, and Avijit Ghosh discussed AI's double-edged nature, a continuum from tool to full-fledged named agent with personality, and the shortcomings of current "benchmark maxing" evaluation; Ghosh proposed an evaluation system that measures both system behavior and impact on people and public participation in reporting and assessment, for example incentivizing companies with liability relief if they fix a reported problem within 60 days.
Harvard Gazette In an NBER working paper, Harvard Kennedy School economists Doug Elmendorf and Karen Dynan with Brookings's Louise Sheiner lay out four scenarios for AI's economic impact, ranging from a moderate GDP boost with little reduction in worker numbers to much faster GDP growth with persistently high unemployment, and estimate that their AI scenario leaves about 3 million people, roughly 2 percent of the labor force, out of work at any given time; separately, Harvard Business School's Joseph Fuller, working with Accenture Research, developed an AI model finding that 41 percent of all work tasks can today be automated or augmented by AI, while only about one-third of firms' AI experiments succeed, which helps explain why mass layoffs have not yet appeared.
In an NBER working paper, Harvard Kennedy School economists Doug Elmendorf and Karen Dynan with Brookings's Louise Sheiner lay out four scenarios for AI's economic impact, ranging from a moderate GDP boost with little reduction in worker numbers to much faster GDP growth with persistently high unemployment, and estimate that their AI scenario leaves about 3 million people, roughly 2 percent of the labor force, out of work at any given time; separately, Harvard Business School's Joseph Fuller, working with Accenture Research, developed an AI model finding that 41 percent of all work tasks can today be automated or augmented by AI, while only about one-third of firms' AI experiments succeed, which helps explain why mass layoffs have not yet appeared.
In an NBER working paper, Harvard Kennedy School economists Doug Elmendorf and Karen Dynan with Brookings's Louise Sheiner lay out four scenarios for AI's economic impact, ranging from a moderate GDP boost with little reduction in worker numbers to much faster GDP growth with persistently high unemployment, and estimate that their AI scenario leaves about 3 million people, roughly 2 percent of the labor force, out of work at any given time; separately, Harvard Business School's Joseph Fuller, working with Accenture Research, developed an AI model finding that 41 percent of all work tasks can today be automated or augmented by AI, while only about one-third of firms' AI experiments succeed, which helps explain why mass layoffs have not yet appeared.
In an NBER working paper, Harvard Kennedy School economists Doug Elmendorf and Karen Dynan with Brookings's Louise Sheiner lay out four scenarios for AI's economic impact, ranging from a moderate GDP boost with little reduction in worker numbers to much faster GDP growth with persistently high unemployment, and estimate that their AI scenario leaves about 3 million people, roughly 2 percent of the labor force, out of work at any given time; separately, Harvard Business School's Joseph Fuller, working with Accenture Research, developed an AI model finding that 41 percent of all work tasks can today be automated or augmented by AI, while only about one-third of firms' AI experiments succeed, which helps explain why mass layoffs have not yet appeared.
Anthropic Anthropic announced a new study in which Anthropic Interviewer, an AI, asks Free, Pro, and Max users of Claude and Claude Code about their positive and negative experiences with AI, what they want AI to change in areas such as work, school, healthcare, and government, and what they want from AI developers; the study runs September 29 to October 6, 2026, takes roughly 15 minutes per interview, and for the first time lets participants choose to make their complete interview and associated country public, with an FAQ explaining the benefits, re-identification risks, and permanence of that choice.
Anthropic announced a new study in which Anthropic Interviewer, an AI, asks Free, Pro, and Max users of Claude and Claude Code about their positive and negative experiences with AI, what they want AI to change in areas such as work, school, healthcare, and government, and what they want from AI developers; the study runs September 29 to October 6, 2026, takes roughly 15 minutes per interview, and for the first time lets participants choose to make their complete interview and associated country public, with an FAQ explaining the benefits, re-identification risks, and permanence of that choice.
Anthropic announced a new study in which Anthropic Interviewer, an AI, asks Free, Pro, and Max users of Claude and Claude Code about their positive and negative experiences with AI, what they want AI to change in areas such as work, school, healthcare, and government, and what they want from AI developers; the study runs September 29 to October 6, 2026, takes roughly 15 minutes per interview, and for the first time lets participants choose to make their complete interview and associated country public, with an FAQ explaining the benefits, re-identification risks, and permanence of that choice.
Anthropic announced a new study in which Anthropic Interviewer, an AI, asks Free, Pro, and Max users of Claude and Claude Code about their positive and negative experiences with AI, what they want AI to change in areas such as work, school, healthcare, and government, and what they want from AI developers; the study runs September 29 to October 6, 2026, takes roughly 15 minutes per interview, and for the first time lets participants choose to make their complete interview and associated country public, with an FAQ explaining the benefits, re-identification risks, and permanence of that choice.
arXiv The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
arXiv The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.
The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.
The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.
The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.
arXiv A Google DeepMind team developed SynthIDBio, which weaves a statistical watermark into both the amino-acid sequence and the 3D shape of AI-designed proteins to mark their machine-generated origin without noticeably compromising function; the team reports that watermarked proteins bound targets involved in viral infection, blood-vessel formation and immune regulation as efficiently as unwatermarked ones, but the tag can in many cases be scrubbed by running a watermarked protein through another design tool, so it is framed as one layer in a layered biosecurity framework rather than a standalone solution.
A Google DeepMind team developed SynthIDBio, which weaves a statistical watermark into both the amino-acid sequence and the 3D shape of AI-designed proteins to mark their machine-generated origin without noticeably compromising function; the team reports that watermarked proteins bound targets involved in viral infection, blood-vessel formation and immune regulation as efficiently as unwatermarked ones, but the tag can in many cases be scrubbed by running a watermarked protein through another design tool, so it is framed as one layer in a layered biosecurity framework rather than a standalone solution.
A Google DeepMind team developed SynthIDBio, which weaves a statistical watermark into both the amino-acid sequence and the 3D shape of AI-designed proteins to mark their machine-generated origin without noticeably compromising function; the team reports that watermarked proteins bound targets involved in viral infection, blood-vessel formation and immune regulation as efficiently as unwatermarked ones, but the tag can in many cases be scrubbed by running a watermarked protein through another design tool, so it is framed as one layer in a layered biosecurity framework rather than a standalone solution.
A Google DeepMind team developed SynthIDBio, which weaves a statistical watermark into both the amino-acid sequence and the 3D shape of AI-designed proteins to mark their machine-generated origin without noticeably compromising function; the team reports that watermarked proteins bound targets involved in viral infection, blood-vessel formation and immune regulation as efficiently as unwatermarked ones, but the tag can in many cases be scrubbed by running a watermarked protein through another design tool, so it is framed as one layer in a layered biosecurity framework rather than a standalone solution.
arXiv SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
arXiv EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
arXiv The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
arXiv The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.
The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.
The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.
The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.
arXiv The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.
The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.
The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.
The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.