AI Core
966 items
Survey reports that AI combining behavioral, genetic, and imaging data with GAN-based augmentation may improve autism spectrum disorder screening, but limited data, class imbalance, and scarce external clinical validation remain
This survey reviews AI, machine learning, deep learning, and generative adversarial network (GAN) approaches to autism spectrum disorder (ASD) screening, focusing on multimodal learning across behavioral, genetic, environmental, neuroimaging, physiological, and clinical data and on the role of GANs in synthetic-data generation and augmentation, concluding that multimodal AI may represent ASD-related characteristics more comprehensively than single-modality approaches while facing challenges of limited and heterogeneous datasets, class imbalance, multimodal integration, GAN training instability, synthetic-data quality, privacy and security, interpretability, generalizability, and limited external clinical validation.
Replacing Euclidean distance with random-forest weights, FORWS and FORWC match or beat conventional tools on noisy high-dimensional series and extract directed interactions between honeybee-hive acoustic vectors and scalar temperature
The study proposes Forest-Weighted S-map (FORWS) and Forest-Weighted Causal Inference (FORWC), which replace the Euclidean metric with adaptive "forest weights" derived from random forest ensembles, and reports comparable or improved forecasting skill relative to conventional tools, substantial resilience to dynamic process noise, mitigation of the curse of dimensionality, and multimodal directed causal inference between high-dimensional acoustic vectors and scalar temperature monitored in a honeybee hive.
A two-branch URL and host-feature model with leakage-resistant OOF stacking detects phishing sites at 92.67% accuracy and 0.9788 ROC-AUC on the UCI dataset
This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
VRPTR predicts individual language activation maps from resting-state fMRI and provides calibrated uncertainty estimates
The study introduces VRPTR, a three-dimensional encoder-decoder combining a compressed Transformer bottleneck, variational latent sampling, and multiscale skip connections, trained on 360 healthy adults from the WU-Minn Human Connectome Project and evaluated on 40 held-out participants for the story-versus-math language contrast, achieving mean voxel-map Pearson r=0.642 and Dice AUC=0.519, exceeding compact volumetric BrainSurfCNN-like and SWIFUN-like comparators by Δr=0.0376 and 0.0335 respectively, while raw 95% intervals covered only 4.7% of observed values and five-fold calibration within the held-out cohort raised coverage to 94.8%.
E-commerce CLV comparison: a neural network reaches RMSE 297.06 on exported scoring outputs, beating three regression and ensemble models
Using 17,049 e-commerce transaction records from 5,000 customers observed between January 2023 and February 2024, the study built a customer-level customer lifetime value (CLV) prediction pipeline with 33 modelling attributes and compared ridge-based Linear Regression, Random Forest, Gradient Boosted Trees and a feed-forward Neural Network under a common 10-fold cross-validation design in Altair AI Studio (RapidMiner), then converted predicted value into three actionable segments of 2,763 low-value, 1,252 medium-value and 985 high-value customers.
Google Research introduces Diffusion Controller: a lightweight steering-damper network that beats LoRA on HPS-v2 win rates in gray-box settings, with a white-box version reaching a 90% win rate over baseline
Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
NVIDIA VSS Blueprint 3.3 builds a visual AI agent from one prompt in under 30 minutes and cuts VLM input tokens by 80%
NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
MaLiang-Harness generates images and video from executable programs: GPT-6-Astra reaches 100% generation success on both benchmarks, with 96.0% of image and 76.9% of video tasks meeting all quality thresholds
The work introduces MaLiang-Harness, a framework that organizes MLLM-driven image and video generation as a persistent process of construction, inspection, and revision, in which a Persistent Executable Generation state, a Traceable Generation Process, and Revision-aware Editing and Verification share a common revision reference; it evaluates 11 and 4 closed-source MLLMs on MaLiang-IBench (50 text-to-image prompts) and MaLiang-VBench (13 text-to-video prompts), measuring GPT-6-Astra at 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
VoxMem tests 15 audio LLMs on 3,196 questions: none tops 40% at 32K, and swapping audio for transcripts drops speaker accuracy from 69.8% to 10.3%
The authors propose a two-axis taxonomy of spoken conversational memory — acoustic evidence type (speech semantics, speaker identity, paralinguistic cues, environmental sound) crossed with memory operation (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal) — and build VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) at four context budgets from 8K to 64K tokens; evaluating 15 large audio language models, no model exceeds 40% overall accuracy at 32K (best 38.5%), models remember what was said far better than who said it, how it was said, or what was audible, and replacing audio with exact transcripts drops speaker accuracy from 69.8% to 10.3% while speech-semantics accuracy barely moves (75.9% to 71.0%).
Page 18 · showing 10