Pretraining on diverse data and fine-tuning with 500 target recordings lifted a sperm whale click-train detector's AUC from as low as 0.60 to 0.78–0.97 on unseen datasets
Synopsis
Starting from a previously trained temporal convolutional network sperm whale click-train detector, the study compared four transfer approaches across four sources (BAL, CS, ICE, MED)—cross-dataset baseline evaluation, training from scratch on only 500 target recordings, pretraining with random fine-tuning, and pretraining with active (uncertainty-based) fine-tuning—and found that pretrained models dropped in performance on unseen data but that fine-tuning with 500 target recordings effectively mitigated the drop, with active fine-tuning consistently outperforming the other approaches although its gain over random fine-tuning was marginal.
Interpretation
Across the four target datasets, pretraining plus fine-tuning (random or active) gave the best performance: median AUC rose from a baseline of 0.79 to 0.95–0.97 for BAL, from 0.60 to 0.80–0.83 for ICE, from 0.71 to 0.78–0.79 for MED, and from 0.81 to 0.88–0.90 for CS. The earlier model from the same team had all six datasets represented in both training and test sets, so transferability to unseen datasets was not assessed; this study restricts evaluation to sources excluded from training. Each experiment used 5-fold cross-validation, reporting median AUC and F1 with standard deviations; the table gives full values for all four target sources.
Training from scratch on only 500 target recordings matched or even exceeded fine-tuning on some datasets (CS: AUC 0.87, F1 0.75, close to random fine-tuning's 0.88/0.80) but was clearly worse on others (BAL: AUC 0.69, F1 0.65; MED: AUC 0.60, F1 0.44). This provides a per-dataset empirical contrast between data volume and data match, showing that the advantage of pretraining depends on how well target data align with training data. Controlled comparison under the same 500-recording budget, with 5-fold cross-validation and both AUC and F1 reported.
Active fine-tuning (uncertainty-based sample selection) consistently outperformed random fine-tuning on all target datasets, but the improvement was marginal; it was considerably worse than random fine-tuning at 150 and 250 fine-tuning samples and only surpassed it at 500. It brings active learning into domain adaptation for PAM detectors and systematically compares four uncertainty sampling strategies—standard, class-balanced, weighted, and class-balanced weighted—along with initial sample size M and per-step batch size N. Four sampling strategies compared against a random baseline across four datasets, with fine-tuning sizes of 150, 250, and 500 and 5-fold cross-validation; the authors also report that standard uncertainty sampling consistently selected more files containing click trains (Figure S3).
Transfer success depended on acoustic match between target and training data: datasets with clearer signals or stricter labeling (CS, BAL) fared better, while datasets with fainter clicks and background noise absent from pretraining (ICE, MED) performed consistently lower under all approaches; qualitative auditing showed the model was often confused by new noise sources such as self-boat noise and other odontocetes. It uses peak-to-median ratio (PMR) to characterize the strongest acoustic events in each 4-min recording and combines this with manual auditing of 10 files per dataset in each of four confidence categories, linking performance differences to dataset acoustics and annotation protocols. PMR is reported separately for target and pretraining datasets (median and standard deviation), alongside an audit table for the four file categories; the authors explicitly note PMR is only a coarse proxy for click strength because most recordings lack click-level labels.
Perspective
The results apply to transferring an existing sperm whale click-train detector to new recording conditions and new study designs, for a binary 4-min-recording task of deciding whether a click train is present. For researchers with a small amount of target-domain annotation, the paper supports pretraining on diverse data and then fine-tuning with about 500 target recordings; when target recordings are consistent and signals are clear (as with CS), training from scratch on a small target sample can also be safe. The authors openly release source code and trained models (github.com/laiagf/WeakTransfer), enabling reproduction and direct reuse.
The dataset used here enforces a per-source 50/50 split between recordings with and without sperm whale clicks, whereas field deployments often face severe class imbalance in which random file selection may be dominated by recordings without the target signal, so whether active fine-tuning still beats random fine-tuning under realistic imbalance remains an open question. The authors also note that active fine-tuning's improvement over random fine-tuning is marginal while adding complexity and accessibility barriers. In addition, PMR is only a coarse proxy for click strength because most recordings lack click-level labels; in the text read here, the body of the results subsections (3.1, 3.2, 3.4) and parts of the figures are incomplete, so the specific numbers and statistical details for those subsections can only be summarized from Tables 1–3 and the discussion.
