ABCP_finder: A Transformer Embedding-Based Prediction of Anti-Breast Cancer Peptides
Synopsis
This work presents ABCP_finder, a computational framework for predicting anti-breast cancer peptides (ABCPs) that combines pretrained protein language model embeddings (ProtBERT and ESM2) with a multilayer perceptron classifier, uses a homology-aware train-test split via CD-HIT at 30% sequence identity with 80% coverage to reduce data leakage, reports ProtBERT as the stronger model with 93.82% accuracy, 86.88% recall, 90.59% F1-score, 0.8618 MCC, 96.67% AUC and a Brier score of 0.0633, selects a 0.7 probability threshold from calibration analysis for high-confidence ABCPs, and shows through external validation with xDeep-AcPEP that unknown peptides predicted as ABCPs exhibit favourable IC values.
Interpretation
It introduces ABCP_finder, described as the first dedicated computational framework for ABCP prediction, representing peptide sequences with pretrained transformer protein language model embeddings (ProtBERT, ESM2) and classifying them with a multilayer perceptron. Where identifying novel ABCPs through experimental methods alone is time consuming and expensive, this work formulates the task as supervised classification over language model embeddings and targets ABCPs specifically. The text describes the full pipeline: construction of positive and negative datasets, embeddings from two pretrained models, and MLP classification, together with reported performance figures, making it a self-consistent methodological report.
Among the tested models ProtBERT performed better, reaching 93.82% accuracy, 86.88% recall, 90.59% F1-score, 0.8618 MCC, 96.67% AUC and a Brier score of 0.0633 under imbalanced conditions. The results provide a comparable head-to-head view of ProtBERT and ESM2 on the same task and report recall, MCC, AUC and Brier score alongside accuracy. The metric values come directly from the text; reporting MCC and Brier score together indicates that evaluation covers both discrimination and probability calibration under class imbalance.
It uses a homology-aware train-test split with CD-HIT at 30% sequence identity and 80% coverage to prevent data leakage and enable realistic evaluation. Compared with a conventional random split, the homology-aware design builds sequence similarity control into the evaluation, so the reported numbers better reflect generalization to new sequences. The split strategy and its thresholds are stated explicitly, making the evaluation design reproducible.
Calibration analysis supports a 0.7 probability threshold for identifying high-confidence ABCPs, and external validation with xDeep-AcPEP shows that unknown peptides predicted as ABCPs by ABCP_finder exhibit favourable IC values. The work folds probability calibration and threshold selection into the prediction pipeline and adds an independent tool as an external check on the biological relevance of predictions for unknown peptides. The threshold choice is backed by calibration analysis; the external validation is a cross-tool consistency observation, with IC values presented as supporting evidence of biological relevance.
Perspective
The framework is intended for large-scale computational screening of anti-breast cancer peptides, suited to research and early drug discovery pipelines that need to prioritize high-confidence ABCPs among many candidates; its evaluation rests on a self-constructed positive and negative dataset, a homology-aware split at CD-HIT 30% identity and 80% coverage, and a 0.7 probability threshold, so the conclusions apply to prediction ranking under this setting rather than to direct clinical or in vivo efficacy.
External validation uses xDeep-AcPEP IC values as support for biological relevance, but the range of unknown peptides covered and the depth of experimental validation are not elaborated in the text; the composition details of the positive and negative datasets, the performance of models beyond ProtBERT and ESM2, and the stability of the 0.7 threshold across different data distributions are directions a careful reader may continue to watch.
