Skip to main content
Back to timeline
arXivSource publication:

AGGRNet splits medical image features into informative and non-informative via a learnable threshold, reaching 5.48% higher accuracy than HiFuse-Small on Kvasir

Synopsis

The work proposes AGGRNet, which embeds a Feature Extraction and Aggregation (FEA) module and a C2PCA block into a YOLOv11 classification backbone; spatial and channel attention plus a learnable threshold τ split feature maps into informative and non-informative parts that are then aggregated by cross-attention, yielding results above the compared SOTA models on five public datasets, including a 5.48% accuracy gain on Kvasir and a 2.2% gain on LIMUC.

Source-provided article image: AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
Figure 1

Figure 1. Comparison of proposed AGGRNet framework with

· Page 1

Interpretation

It proposes the FEA module, which adds spatial and channel attention outputs, applies a sigmoid to obtain attention scores S, and uses a learnable threshold τ (initialized at 0.5) to build complementary binary masks Winfo and Wninfo that separate informative from non-informative features. Unlike standard feature extraction that treats all spatial regions and channels uniformly, this design explicitly separates diagnostically relevant regions from background and lets the threshold adapt during training. The method is given in full equations (Eqs. 1-7) with τ initialized at 0.5; the ablation shows accuracy rising from 0.791 to 0.793 to 0.807 as FEA modules are added progressively.

It proposes the FAM cross-attention with Q = Xinfo + Xninfo, K = Xinfo − Xninfo, and V = Xinfo, biasing attention toward informative features. The authors expand QKT into ‖Xinfo‖² − ‖Xninfo‖² plus cross terms, arguing the positive bias term promotes informative regions while the negative term suppresses non-informative ones, forming a contrast-based cross-attention. The derivation appears in Eq. 11; aggregation uses the softmax attention of Eqs. 12-13, and Eq. 14 adds a residual connection that retains the original features.

It replaces the C2PSA self-attention block in the YOLOv11 classification backbone with the C2PCA block based on channel attention. The stated rationale is that FEA already handles spatial and cross-feature attention, so C2PCA can focus on channel-wise prioritization, is computationally more efficient for final-layer processing, and suits medical images where different channels encode different anatomical or pathological patterns. The ablation shows accuracy rising from 0.775 to 0.793, i.e. +2%; the channel expansion, branch split, channel attention, two residual connections, and concatenation are specified in Algorithm 1.

It reports results above the compared SOTA models on five public datasets: Kvasir accuracy 0.916 vs HiFuse-Small 0.8612 (+5.48%), LIMUC 0.810 vs CDW-CE (Inception-v3) 0.788 (+2.2%), ISIC2018 0.871 vs HiFuse-Base 0.8585 (+1.25%), PathMNIST 0.926 vs ResNet-50 (28) 0.911 (+1.5%), and RetinaMNIST 0.537 vs Google AutoML Vision 0.531 (+0.6%). One framework covers both disease-subtype classification and severity grading, whereas the compared CDW-CE and SATOMIL approaches are confined to ordinal-loss or patient-level multiple-instance settings respectively. Results are tabulated with accuracy, Macro-F1, QWK, MAE, precision, recall, and AUC; AGGRNet has 38.65M parameters versus 127.80M for HiFuse-Base; training uses ImageNet-pretrained weights, 224×224 inputs, and SGD (lr 0.01, momentum 0.937, weight decay 5×10⁻⁴).

Perspective

The work targets two medical image tasks, disease-subtype classification and severity grading, in single-image classification settings on public datasets spanning endoscopy (LIMUC, Kvasir), dermoscopy (ISIC2018), histopathology (PathMNIST), and retinal imaging (RetinaMNIST). FEA is designed to be architecture-agnostic, insertable into deeper CNN layers with a residual connection, so it can be tried on other backbones or medical image classification tasks. For a reader, this means the paper offers a reusable feature-separation and aggregation component plus comparable benchmark results, not a diagnostic tool already deployed in clinical workflows.

A careful reader would still watch: the convergence behavior and final value of the learnable threshold τ are not reported in the main text; the choice of FEA insertion positions rests on the authors' analysis of CNN feature hierarchies, and the optimal placement under other backbones remains to be explored; all results are offline evaluations on public datasets, without prospective clinical validation or inter-reader agreement assessment; the text mentions additional experiments in the supplement, but that part is not included in the loaded text, so its specific content cannot be summarized here.

Sources