Skip to main content
Back to timeline
Natural Sciences and Applied TechnologySource publication:

A two-branch URL and host-feature model with leakage-resistant OOF stacking detects phishing sites at 92.67% accuracy and 0.9788 ROC-AUC on the UCI dataset

Synopsis

This study proposes a hybrid machine learning phishing detection system with two branches trained separately on URL and host-based feature sets and combined through a leakage-resistant Out-Of-Fold stacking mechanism; on the UCI Phishing Websites Dataset with 11,055 instances and 17 selected features it achieves 92.67% accuracy and 0.9788 ROC-AUC, outperforming single feature-based models, with cross-validation showing low-variance generalization and feature importance analysis indicating that interaction-based meta-features strongly influence classification accuracy.

AI-generated editorial illustration: A Hybrid Machine Learning Framework for Phishing Website Detection Using URL and Host-Based Features with Leakage-Resistant Stacking

Interpretation

It introduces a two-branch architecture that trains separate classifiers on URL and host feature sets and fuses them via a leakage-resistant Out-Of-Fold stacking mechanism that derives meta-features from unseen data. Compared with prior approaches relying on a single feature set or on blacklist and rule-based methods, this design models two complementary feature groups separately before fusion and emphasizes that meta-features come from unseen data to avoid data leakage. The paper description explicitly states the two-branch structure and the OOF stacking mechanism, and evaluates the model on the UCI Phishing Websites Dataset with 11,055 instances and 17 features.

On the UCI Phishing Websites Dataset, the hybrid model reaches 92.67% accuracy and 0.9788 ROC-AUC, outperforming models based on a single feature set. The result quantifies the gain from two-branch fusion over single-feature baselines using both accuracy and ROC-AUC. The numbers come from the paper's reported experiments, and cross-validation shows low variance, indicating relatively stable generalization.

It adds interaction-based meta-features, and feature importance analysis shows these meta-features are highly influential in improving classification accuracy. Beyond conventional stacking meta-features, interaction terms supply additional discriminative information, and feature importance analysis locates their contribution. The paper reports feature importance analysis indicating a notable role for interaction meta-features in accuracy improvement.

The model balances performance, computational efficiency, and interpretability, and its leakage-resistant design supports unbiased evaluation and real-world cybersecurity applicability. Rather than pursuing accuracy alone, the work folds methodological rigor (avoiding data leakage) together with efficiency and interpretability as design goals. The paper supports this combined positioning with accuracy, ROC-AUC, cross-validation variance, and feature importance evaluations.

Perspective

The result targets research and engineering settings that classify phishing websites using URL and host-based features, and applies to detection pipelines with the corresponding feature extraction capability; the leakage-resistant OOF stacking design makes evaluation closer to unbiased and can serve as a reference baseline for real cybersecurity deployment. For researchers wishing to reproduce or extend the two-branch fusion idea, the work provides a reusable architecture and evaluation paradigm.

Because the loaded text is an incomplete description rather than the full paper, it is not possible to confirm the specific classifier types used in each branch, hyperparameter settings, the number of cross-validation folds, or full comparison details with other methods; the exact construction of interaction meta-features and their importance values are also not given in the text. These are open questions to watch for when reading the original paper.

Sources