AI Core
961 items
SCITFS folds adaptive redundancy penalization and bootstrap stability regularization into one objective, lifting SVM accuracy by 3.7% and Random Forest by 4.2% on eight benchmarks while cutting 92.3% of features.
The work proposes Stability-Constrained Information-Theoretic Feature Selection (SCITFS), which integrates conditional entropy, normalized mutual information maximization, and a stability term penalizing feature-ranking variance across bootstrap samples into a single formally defined objective, implemented via a greedy forward-selection strategy with proven monotonicity guarantees at O(B n^2 m d^4 + k n^2); across eight benchmark datasets and five classifiers (SVM, Random Forest, k-NN, XGBoost, Logistic Regression), SCITFS outperforms Information Gain, Mutual Information, mRMR, ReliefF, Fisher Score, and JMI, achieving a 3.7% average accuracy improvement on SVM and 4.2% on Random Forest with 92.3% feature reduction while maintaining performance; Friedman test (χ² = 127.4, p < 0.
Evo 2 shows in-context learning on five binary classification tasks with F1 up to 0.902 on short sequences, but collapses at kilobase scale and the 7B model beats the 40B
This study maps the in-context learning operating regime of Evo 2, a nucleotide-level foundation genomic language model, across five binary classification tasks spanning biological and artificial sequences, finding robust performance on shorter natural sequences (F1=0.902 for miRNA, 0.785 for Toxins), degradation with sequence length and collapse at kilobase scale, no benefit from model scaling (the 7B model systematically outperforms the 40B variant), poor prediction of accuracy by perplexity, and mechanistic interpretability via logit-lens and Jacobian Scope suggesting a prediction-generalisation trade-off and that models might track prompt structure rather than signal-carrying content.
Ridezy derives driver credibility in real time from edge AI and IoT sensors and anchors hashes and reputation updates on Polygon, outperforming rating-based, AI-only, and blockchain-only baselines in behavioural fidelity and trust guarantees
The paper presents Ridezy, a decentralized trust architecture that uses edge Artificial Intelligence and Internet of Things (AIIoT) to continuously monitor behavioural indicators such as lane discipline, speed compliance, braking behaviour, and traffic sign adherence, processes behavioural summaries off-chain while anchoring only cryptographic hashes, credibility updates, and payment records on Polygon smart contracts, and compares it against three representative trust models—a traditional rating-based system, an AI-only architecture, and a blockchain-only architecture—across behavioural detection performance, end-to-end latency, cost efficiency, throughput, reputation stability, tamper resistance, and component-wise ablation studies, with results showing higher behavioural fidelity and st
Gated-attention multiple instance learning triages multi-center cervical cytology slides without per-cell labels, reaching 90.96% accuracy with MobileNetV2 and 80.43% balanced accuracy out-of-distribution with Xception
The work presents a weakly-supervised, detection-free multiple instance learning framework that uses dual-branch gated attention pooling to make slide-level predictions on whole-slide cervical cytology images, treating each slide as a bag of local instance patches so that single-cell bounding boxes or pixel-level annotations are not required; evaluated on internal multi-center cohorts (SIPaKMeD, Herlev, and CRIC) and on the unannotated, out-of-distribution Mendeley LBC validation cohort processed via an unsupervised marker-controlled watershed pipeline, it reports that the lightweight MobileNetV2 backbone optimizes in-distribution multi-center accuracy (90.96% accuracy, 0.9800 ROC-AUC) while the higher-capacity Xception provides better out-of-distribution robustness under domain shift (80.
Fuzzy-rank feature selection plus H2O AutoML ensembles reach up to 95.1% accuracy and 98.1% AUC on two public cervical cancer datasets
The work introduces an interpretable Fuzzy Rank-H2O AutoML framework in which a Fuzzy Rank Feature Selection (FRFS) algorithm picks predictors by combining statistical significance, information gain, clinical importance, and uncertainty, H2O AutoML then automatically builds ensemble models, and SHAP and LIME supply global and patient-level explanations; evaluated on two public cervical cancer datasets with stratified five-fold cross-validation where SMOTE is applied only to training folds to avoid information leakage, it reports a highest accuracy of 95.1% and AUC of 98.1%, outperforming traditional machine learning models and the baseline H2O AutoML framework.
Survey reports that AI combining behavioral, genetic, and imaging data with GAN-based augmentation may improve autism spectrum disorder screening, but limited data, class imbalance, and scarce external clinical validation remain
This survey reviews AI, machine learning, deep learning, and generative adversarial network (GAN) approaches to autism spectrum disorder (ASD) screening, focusing on multimodal learning across behavioral, genetic, environmental, neuroimaging, physiological, and clinical data and on the role of GANs in synthetic-data generation and augmentation, concluding that multimodal AI may represent ASD-related characteristics more comprehensively than single-modality approaches while facing challenges of limited and heterogeneous datasets, class imbalance, multimodal integration, GAN training instability, synthetic-data quality, privacy and security, interpretability, generalizability, and limited external clinical validation.
Replacing Euclidean distance with random-forest weights, FORWS and FORWC match or beat conventional tools on noisy high-dimensional series and extract directed interactions between honeybee-hive acoustic vectors and scalar temperature
The study proposes Forest-Weighted S-map (FORWS) and Forest-Weighted Causal Inference (FORWC), which replace the Euclidean metric with adaptive "forest weights" derived from random forest ensembles, and reports comparable or improved forecasting skill relative to conventional tools, substantial resilience to dynamic process noise, mitigation of the curse of dimensionality, and multimodal directed causal inference between high-dimensional acoustic vectors and scalar temperature monitored in a honeybee hive.
A two-branch URL and host-feature model with leakage-resistant OOF stacking detects phishing sites at 92.67% accuracy and 0.9788 ROC-AUC on the UCI dataset
This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.
VRPTR predicts individual language activation maps from resting-state fMRI and provides calibrated uncertainty estimates
The study introduces VRPTR, a three-dimensional encoder-decoder combining a compressed Transformer bottleneck, variational latent sampling, and multiscale skip connections, trained on 360 healthy adults from the WU-Minn Human Connectome Project and evaluated on 40 held-out participants for the story-versus-math language contrast, achieving mean voxel-map Pearson r=0.642 and Dice AUC=0.519, exceeding compact volumetric BrainSurfCNN-like and SWIFUN-like comparators by Δr=0.0376 and 0.0335 respectively, while raw 95% intervals covered only 4.7% of observed values and five-fold calibration within the held-out cohort raised coverage to 94.8%.
Page 17 · showing 10