Skip to main content

Daily report

AI and science frontiers · 2026-09-13

Only content delivered through the publication boundary on this date is included.

medRxiv

A Large Language Model for Risk-of-Bias Assessment in Systematic Reviews of Prognosis Studies in Clinical Neurology

This study designed a zero-shot prompted LLM pipeline as a virtual mimic of a human reviewer for the QUIPS framework and applied it to 298 articles from previously published systematic reviews of prognosis studies in neurology across three domains (epilepsy, traumatic brain injury, stroke), finding limited LLM-human agreement (weighted kappa = 0.22, 95% CI 0.12-0.33) while tentatively suggesting, based on a small sample (n=5), that it may not be inferior to human-human agreement (weighted kappa = -0.25, 95% CI -1.04-0.54), with Wilcoxon signed-rank tests statistically significant (p < 0.05) across four bias domains and overall risk scores, and rank-biserial correlations demonstrating human raters' tendency to assign higher risk scores than LLM counterparts.
bioRxiv

Agentic-AI-ready genome-wide poxvirus-host interaction screen refined by a protein language model

This work proposes ICARus, a positive-unlabelled read-out refinement framework that integrates protein-protein interaction information derived from a protein language model into a genome-wide RNA interference screen to boost the discovery of vaccinia virus-host interactions and human genes with potential antiviral function, and releases the raw and refined read-outs of the screen as an agentic-AI-enabled community resource.
bioRxiv

Iterative Gene Enrichment Analysis: interpretable networks for human and AI-assisted biological insights

This work introduces iterative Gene Enrichment Analysis (iGEA), a software framework that repeatedly selects the most significant enriched term and removes its overlapping genes to yield compact non-overlapping enriched terms within each gene-set collection, then integrates collection-specific results into a cross-collection gene-term network; applied to a published set of HIV dependency factors, iGEA identified five compact modules spanning secretory trafficking, nuclear transport, transcription elongation, proteostasis, and innate immune signaling, providing a structured representation for human interpretation and LLM-assisted exploration.
bioRxiv

ABCP_finder: A Transformer Embedding-Based Prediction of Anti-Breast Cancer Peptides

This work presents ABCP_finder, a computational framework for predicting anti-breast cancer peptides (ABCPs) that combines pretrained protein language model embeddings (ProtBERT and ESM2) with a multilayer perceptron classifier, uses a homology-aware train-test split via CD-HIT at 30% sequence identity with 80% coverage to reduce data leakage, reports ProtBERT as the stronger model with 93.82% accuracy, 86.88% recall, 90.59% F1-score, 0.8618 MCC, 96.67% AUC and a Brier score of 0.0633, selects a 0.7 probability threshold from calibration analysis for high-confidence ABCPs, and shows through external validation with xDeep-AcPEP that unknown peptides predicted as ABCPs exhibit favourable IC values.
bioRxiv

Large Language Models Predict Human Social Behavior via Interpretable Mechanisms

This study introduces MindEvolve, an autonomous workflow in which multiple large language models generate interpretable symbolic models of cognition for a battery of socioeconomic games spanning four core domains of social cognition—economic preferences, social preferences, social reasoning (theory of mind), and recursive planning—with expert human evaluators assessing interpretability and theoretical coherence; results show that most LLMs robustly capture economic and social preferences in relatively simple strategic settings and in some cases generate novel models integrating broader knowledge than those proposed by human experts, while their capacity to model more complex psychological processes remains limited, though a subset of state-of-the-art models shows promising performance on h
AL-IMAM Journal on Islamic Studies Civilization and Learning Societies

Exploitation of Public Figures' Faces by AI Platforms: Unjust Enrichment and Layered Accountability under Indonesian Law

Using normative legal research grounded in the Copyright Law, the Electronic Information and Transactions Law, the Personal Data Protection Law, and Article 1365 of the Civil Code, this study argues that a public figure's face has shifted within generative AI ecosystems from a marker of identity into an intangible economic asset sustaining platform valuation, while existing legal regimes remain fragmented rather than complementary—copyright refuses to treat a face as a protectable work and data protection law leans toward privacy rather than economic interest; this misalignment allows the four elements of unjust enrichment (enrichment, loss, causation, and absence of legal basis) to be satisfied and opens the possibility of qualifying the conduct as unlawful under Article 1365, on which ba
bioRxiv

OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation

The study presents OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human immune-sequence records that uses sequence-type tokens and complementary component orders to jointly learn from individual TCR chains and partial or complete TCR-pMHC associations, thereby performing both TCR recognition prediction and conditional sequence generation within one model, achieving AUPRCs of 0.7009 for peptide-TCRβ recognition and 0.8235 for TCR-pMHC interaction prediction on unseen epitopes and a mean AUROC of 0.9436 in distinguishing cancer from healthy repertoires across 11 independent pan-cancer cohorts.
medRxiv

Adolescent engagement and sentiment toward reproductive-health videos on Chinese social media: a cross-sectional content analysis

Using an 18-keyword query, this study retrieved 743 Bilibili reproductive-health videos published between September 2016 and June 2024, covering 486.8 million views, 12.3 million likes, and roughly 760,000 textual entries, described engagement metrics with descriptive and Pearson-correlation analyses, and applied a three-class sentiment classifier powered by Qwen2.5-32B to comments and danmaku, finding that adolescent engagement centres on basic contraception while niche or contentious topics (sterilisation surgery 3.0%, fertility-awareness-based methods 2.3%) attract disproportionately high interaction, only 17% of uploaders had confirmed medical training, and negative sentiment predominated for induced abortion (52.2%) and novel contraceptives (52.
medRxiv

Hybrid lexical-semantic retrieval over SNOMED CT: combining two retrieval paradigms to facilitate clinical data entry

The work proposes and implements a hybrid retrieval architecture that lets deterministic lexical matching and learned semantic matching coexist over SNOMED CT's own curated descriptions, combining multi-prefix search, BioLORD-2023-M embeddings, an optional BGE cross-encoder for re-ranking, and a local LLM for query normalization, with Reciprocal Rank Fusion and a hierarchy filter; on a search-only linking evaluation over 542 disease mentions from the DisTEMIST corpus (Spanish, zero-shot), semantic search with re-ranking and no LLM query pre-processing reached accuracy@1 of 0.60 and recall@10 of 0.80, and on 12,897 mentions from English real-EHR discharge notes a field-scoped typeahead placed the concept on the top-10 picker list for 73% of mentions (accuracy@1 0.