Medicine & Health
469 items
OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation
The study presents OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human immune-sequence records that uses sequence-type tokens and complementary component orders to jointly learn from individual TCR chains and partial or complete TCR-pMHC associations, thereby performing both TCR recognition prediction and conditional sequence generation within one model, achieving AUPRCs of 0.7009 for peptide-TCRβ recognition and 0.8235 for TCR-pMHC interaction prediction on unseen epitopes and a mean AUROC of 0.9436 in distinguishing cancer from healthy repertoires across 11 independent pan-cancer cohorts.
Agentic-AI-ready genome-wide poxvirus-host interaction screen refined by a protein language model
This work proposes ICARus, a positive-unlabelled read-out refinement framework that integrates protein-protein interaction information derived from a protein language model into a genome-wide RNA interference screen to boost the discovery of vaccinia virus-host interactions and human genes with potential antiviral function, and releases the raw and refined read-outs of the screen as an agentic-AI-enabled community resource.
ABCP_finder: A Transformer Embedding-Based Prediction of Anti-Breast Cancer Peptides
This work presents ABCP_finder, a computational framework for predicting anti-breast cancer peptides (ABCPs) that combines pretrained protein language model embeddings (ProtBERT and ESM2) with a multilayer perceptron classifier, uses a homology-aware train-test split via CD-HIT at 30% sequence identity with 80% coverage to reduce data leakage, reports ProtBERT as the stronger model with 93.82% accuracy, 86.88% recall, 90.59% F1-score, 0.8618 MCC, 96.67% AUC and a Brier score of 0.0633, selects a 0.7 probability threshold from calibration analysis for high-confidence ABCPs, and shows through external validation with xDeep-AcPEP that unknown peptides predicted as ABCPs exhibit favourable IC values.
Adolescent engagement and sentiment toward reproductive-health videos on Chinese social media: a cross-sectional content analysis
Using an 18-keyword query, this study retrieved 743 Bilibili reproductive-health videos published between September 2016 and June 2024, covering 486.8 million views, 12.3 million likes, and roughly 760,000 textual entries, described engagement metrics with descriptive and Pearson-correlation analyses, and applied a three-class sentiment classifier powered by Qwen2.5-32B to comments and danmaku, finding that adolescent engagement centres on basic contraception while niche or contentious topics (sterilisation surgery 3.0%, fertility-awareness-based methods 2.3%) attract disproportionately high interaction, only 17% of uploaders had confirmed medical training, and negative sentiment predominated for induced abortion (52.2%) and novel contraceptives (52.
Arti-JEPA: Adapting a Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
The work continues the self-supervised objective of the general video world model V-JEPA 2 on roughly 62 hours of unlabelled vocal-tract real-time MRI to obtain a frozen Arti-JEPA encoder, and evaluates it on phoneme prediction, fluent-versus-disfluent classification, and pre/post-glossectomy transfer, finding that a temporal video prior clearly beats per-frame image encoders, that latent prediction is at least as strong as pixel reconstruction, that domain adaptation roughly doubles cross-domain Cohen's kappa for phonemes but helps binary stuttering detection only marginally, and that phoneme signal remains partly decodable after surgery.
HierSTT: A Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting
The work proposes HierSTT, an end-to-end hierarchical Transformer framework that uses a Temporal Fusion Transformer for national-level ED demand and spatio-temporal Transformer encoder-decoders to produce regional and hospital forecasts conditioned on higher-level predictions, with a coherence-aware loss penalizing cross-level inconsistency; it also releases a nationwide Portuguese ED dataset covering 81 hospitals across 5 regional health administrations with heterogeneous level-specific covariates, and reports a 32% average WAPE reduction versus the best non-hierarchical deep learning baseline, outperforming classical hierarchical reconciliation methods while producing near-coherent predictions across levels.
Sparse Autoencoders Localize the Clinical Triage Format Effect: Medical Features Go Inactive at the Decision Token, Scaffold Features Account for Over 91% of Gemma Attribution
Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B on clinical triage vignettes, this study finds that medical features fire on the shared clinical narrative under both formats but are inactive at the multiple-choice decision token; emergency-tier information is linearly decodable from vignette representations (ROC-AUC 0.95–1.00) with no significant format difference yet is attenuated at the decision token, and scaffold-peaking features account for over 91% of unsigned attribution in both Gemma models, placing the strongest correlates of the format effect at answer selection.
From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation
The work proposes a generation-oriented 3D-CLIP encoder trained with structured hard negatives constructed exclusively at the text level (Attribute-Aware Negatives, AAN, and Semantic-Aware Negatives, SAN) to strengthen contrastive learning under the small-batch constraints of volumetric encoders, and uses it to condition a fully end-to-end latent diffusion model operating directly in 3D latent space, achieving lower FID, higher pathology-classification AUC and precision, faster inference, and lower GPU memory than competing methods on CT-RATE across 18 pathological conditions.
All results are shown
Page 47 · showing 9