SCITFS folds adaptive redundancy penalization and bootstrap stability regularization into one objective, lifting SVM accuracy by 3.7% and Random Forest by 4.2% on eight benchmarks while cutting 92.3% of features.
Synopsis
The work proposes Stability-Constrained Information-Theoretic Feature Selection (SCITFS), which integrates conditional entropy, normalized mutual information maximization, and a stability term penalizing feature-ranking variance across bootstrap samples into a single formally defined objective, implemented via a greedy forward-selection strategy with proven monotonicity guarantees at O(B n^2 m d^4 + k n^2); across eight benchmark datasets and five classifiers (SVM, Random Forest, k-NN, XGBoost, Logistic Regression), SCITFS outperforms Information Gain, Mutual Information, mRMR, ReliefF, Fisher Score, and JMI, achieving a 3.7% average accuracy improvement on SVM and 4.2% on Random Forest with 92.3% feature reduction while maintaining performance; Friedman test (χ² = 127.4, p < 0.
Interpretation
SCITFS places adaptive redundancy penalization and bootstrap-based stability regularization inside one formally defined objective, paired with a greedy forward-selection strategy with proven monotonicity guarantees and complexity O(B n^2 m d^4 + k n^2). Traditional information-theoretic approaches such as Mutual Information and mRMR typically separate relevance ranking from stability handling; this work writes the stability term directly into the objective so ranking variance becomes an optimizable component. The text provides a formal objective definition, monotonicity guarantees, and a complexity expression, which is a method-level argument; the abstract does not include the derivation details.
On eight benchmark datasets and five classifiers, SCITFS achieves higher accuracy than Information Gain, Mutual Information, mRMR, ReliefF, Fisher Score, and JMI, with a 3.7% average gain on SVM and 4.2% on Random Forest, while reducing features by 92.3% with maintained performance. Relative to existing information-theoretic and filter baselines, the work reports both cross-classifier accuracy gains and large feature reduction rather than comparing on a single metric. Based on experiments over eight benchmark datasets and five classifiers, with statistical validation via the Friedman test (χ² = 127.4, p < 0.001) and the Wilcoxon signed-rank test with Bonferroni correction.
Across 100 bootstrap trials, SCITFS yields 89.6% feature-ranking consistency versus 67.4% for mRMR, corroborated by a supplementary comparison against a Jensen–Shannon-divergence-based stability metric. The work treats ranking stability as a quantifiable evaluation dimension and cross-checks it against another stability metric, rather than reporting predictive accuracy alone. The stability analysis rests on 100 bootstrap trials plus a supplementary comparison against a Jensen–Shannon-divergence-based stability metric.
An ablation study isolates the contribution of each objective term, and SCITFS remains robust to feature noise up to 20% contamination. The ablation attributes overall gains to specific objective terms and additionally characterizes behavior under noise, going beyond end-to-end performance reporting. The ablation and noise-contamination experiments are described conclusively in the abstract; specific values are not expanded in the text.
Perspective
The result targets high-dimensional supervised learning settings that need stable feature selection under noise and redundancy, and it applies to tabular-feature tasks such as bioinformatics, genomics, cybersecurity, finance, and high-dimensional pattern recognition. For practitioners, it offers a route to write stability directly into the objective function, making feature rankings more reproducible across data partitions and substantially compressing feature counts while maintaining performance. The method is implemented via greedy forward selection with complexity O(B n^2 m d^4 + k n^2), so its applicability depends on the relative scale of the number of bootstrap rounds B, sample size n, feature count m, and number of selected features k.
The currently visible text is an incomplete abstract-style description; it does not include specific dataset names, per-baseline numerical results, the individual contribution size of each ablated objective term, the full noise-experiment curves, or behavior beyond 20% contamination. Readers interested in those details would need the tables and figures of the full paper. In addition, the stability advantage is currently compared mainly against mRMR and one Jensen–Shannon-divergence-based stability metric, so whether other stability-oriented methods show a similar pattern remains an open question to watch.
