Skip to main content
Back to timeline
arXivSource publication:

Compact CNN reaches 0.7226 macro F1 on DeepShip under recording-level partitioning, while ResNet18 yields no validation gain

Synopsis

The work proposes a compact underwater acoustic classification framework combining multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations; it first evaluates lightweight classifiers and CNNs on the ShipsEar dataset (a two-layer CNN reaches 0.9918 macro F1 and an RBF-SVM reaches 0.9883), but source-recording provenance cannot be reconstructed, preventing verification of recording-independent generalisation, so it then evaluates on the DeepShip dataset using recording-level partitioning before segmentation, where a 157K-parameter compact CNN achieves a test macro F1 of 0.7226 while an 11.

Source-provided article image: Towards Deployable Underwater Vessel Classification
Figure 1 ·

Figure 1: Example ShipsEar waveform and corresponding conventional and auditory-inspired acoustic representations used in this study.

arXiv

Interpretation

It proposes a compact underwater acoustic classification framework for deployability, integrating multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations. Relative to single-representation or single-model conventions, the framework examines multiple conventional and auditory-inspired representations together with compact convolutional structures within one design pipeline. Supported by the framework description and experimental evaluation on the ShipsEar and DeepShip datasets, taking the form of a method proposal plus experimental validation.

On the provided ShipsEar split, a two-layer CNN achieves a macro F1 of 0.9918 and an RBF-SVM reaches 0.9883. It provides a concrete performance comparison between lightweight classifiers and CNNs under that dataset split. The numbers come from the reported evaluation on the provided ShipsEar split; however, the authors note that source-recording provenance cannot be reconstructed, so these high scores cannot verify recording-independent generalisation.

On the DeepShip dataset, using recording-level partitioning before segmentation, a 157K-parameter compact CNN achieves a test macro F1 of 0.7226, while an 11.17M-parameter ResNet18 provides no improvement in validation performance under the matched setting. It moves evaluation from segment-level splitting to recording-level partitioning and compares a compact model against a larger model under this stricter protocol. The results come from the reported DeepShip recording-level partitioning experiments, including specific parameter counts, macro F1, and the ResNet18 validation comparison.

The results demonstrate that representation-aware feature and model design, together with rigorous recording-level evaluation, matter for both classification performance and deployability in compact underwater acoustic systems. It attributes performance differences to how well representations and model design match, and to the choice of evaluation protocol, rather than to model size alone. The conclusion is supported jointly by the ShipsEar and DeepShip experiments, with the DeepShip recording-level protocol providing evidence closer to deployment conditions.

Perspective

The framework targets underwater acoustic classification scenarios that need deployment under resource-constrained conditions, and is suited to researchers and practitioners concerned with model size and recording-level evaluation protocols. Its conclusions rest on the ShipsEar and DeepShip datasets and on the conventional and auditory-inspired representations examined in the text, with the recording-level partitioning protocol being the key premise for trusting the results.

The high macro F1 on ShipsEar cannot be confirmed as recording-independent because source-recording provenance cannot be reconstructed, a point the text states explicitly, so readers should treat it as an open question rather than directly citable generalisation evidence. For DeepShip, the 0.7226 macro F1 and the finding that ResNet18 shows no validation improvement are reported without the training details, representation combinations, and hyperparameter settings in this summary, so reproducibility requires consulting the original. In addition, how the compact model performs in other sea areas, on other vessel types, or under different signal-to-noise conditions remains an open question.

Sources