Skip to main content
Back to timeline
Nature CommunicationsSource publication:

Topology-Enhanced Machine Learning for Speech Signal Processing

Synopsis

This work introduces TopCap and TopNN, which use time-delay embedding and persistent homology to extract topological features (such as maximal persistence and its birth time) from speech time series; on voiced versus voiceless consonant classification, TopCap reaches accuracy comparable to some state-of-the-art neural networks on small datasets while offering greater efficiency and interpretability, and TopNN, which concatenates topological features with gated recurrent unit features, achieves higher accuracy, steadier performance, and stronger noise robustness than standard neural networks across multiple datasets and signal-to-noise ratios.

Source-provided article image: Topology-enhanced machine learning for speech signal processing
Figure 1 ·

Figure 1: Visualisation for data and some of their commonly used topological descriptors. a Time-delay embedding (dimension=3, delay=10, skip=1) of f ⁡ ( t n ) = sin ⁡ ( 2 ​ t n ) − 3 ​ sin ⁡ ( t n ) f(t_{n})=\sin(2t_{n})-3\sin(t_{n}) , with t n = π 50 ​ n t_{n}=\frac{\pi}{50}n ( 0 ≤ n ≤ 200 0\leq n\leq 200 ). Resulting point clouds lay on a closed curve in 3-dimensional Euclidean space. The colour indicates original locations of data in the time series. b A topological space and its triangulation. On the left is a topological space consisting of a 1-dimensional sphere (i.e., a circle) and a 2-dimensional sphere with a single point of contact, denoted as 𝕊 1 ∨ 𝕊 2 \mathbb{S}^{1}\vee\mathbb{S}^{2} . The right depicts a triangulation of this topological space. c Average temperature in the U.S. with monthly values (dark blue dots) and yearly values (green curve). The left panel shows a single-year section of average temperature. d Computing PH. The four plots consecutively show how a persistence diagram or barcode is computed: Connect each pair of points with a distance less than ϵ \epsilon by a line segment, fill in each triple of points with mutual distances less than ϵ \epsilon with a triangular region, etc., and compute the corresponding homology groups. e Characterising the vibration of a time series in terms of its variability of frequency, amplitude, and average line. f Commonly used representations for PH, with an example of 100 points uniformly distributed over a bounded region in 2D Euclidean space. A persistence barcode is a multiset of intervals, where the horizontal axis shows when each feature appears and disappears. A persistence diagram directly plots the birth and death times (values of ϵ \epsilon , not to be confused with time series) of each interval. In both plots, 0 and 1 correspond to the 0-dimensional loops (connected components) and 1-dimensional loops. In a persistence landscape, the k k th landscape is the k k th largest value of tent functions for each feature, here taken as the 1-dimensional loops, with the horizontal axis representing resolution (turning the persistence diagram clockwise by 45 ∘ 45^{\circ} ). Similarly, a persistence image is created by applying Gaussian functions centred at each feature and then converting them into a pixelated image, where both the horizontal and vertical axes represent resolution.

arXiv

Interpretation

Proposes TopCap: time-delay embedding (dimension d=100, delay τ=6T/d) combined with persistent homology, extracting maximal persistence and its birth time from the 1-dimensional persistence diagram as features fed to Tree, Discriminant, Regression, Naive Bayes, SVM, k-NN, Linear, and Ensemble classifiers. Unlike prior speech methods relying on energy and spectral information (STFT, MFCC), this pipeline directly characterizes the topological structure of the time series, and just two quantities, maximal persistence and birth time, suffice to separate voiced from voiceless consonants. On 5101 records (3571 training, 1530 test; 2138 voiced and 2963 voiceless), most algorithms exceed 96% AUC and 93% accuracy; voiced consonants tend to show higher birth time and lifetime.

Benchmarks TopCap against STFT–CNN, MFCC–GRU, and MFCC–Transformer on 8 small and 4 large datasets. TopCap exceeds MFCC–GRU and STFT–CNN-8 on small datasets but generally falls short of deep neural networks on large datasets, a gap that motivates integrating topological features into neural networks. Small-dataset accuracies include 94.3% on HT1, 94.6% on LJ, and 83.9% on TIMIT; large-dataset accuracies include 92.5% on ALLSSTAR, 92.9% on LJSpeech, 92.8% on TIMIT, and 88.7% on LibriSpeech.

Proposes TopNN: concatenating TopCap's maximal persistence feature with the GRU's 6-dimensional hidden state into a 7-dimensional vector passed through a fully connected decoder, with ZeroNN (topological feature replaced by zeros) as a control. This design attributes performance differences to the inclusion of topological features rather than model capacity; TopNN outperforms ZeroNN and NN on both clean and noisy data, with the gap widening as noise increases. On ALLSSTAR, LJSpeech, and TIMIT, TopNN achieves higher training and test accuracy from no noise through strong noise (SNR=0dB), with lower accuracy variance across repeated experiments.

Explores on synthetic and real data how persistence diagrams characterize non-periodic vibration patterns, and proposes formant spectral features and cyclic time-delay embedding configuration eigenvalues as additional geometric features. Synthetic experiments show that point distributions in the lower region of the diagram distinguish three fundamental variations of frequency, amplitude, and average line; in real speech, unstable series show higher point density in the lower region while stable series tend to attain high maximal persistence. Synthetic data use dimension 100, delay 3, skip 10; real data come from two vowel [A] recordings in ALLSSTAR cropped into four overlapping intervals ending at 600, 800, 1000, and 1200; the 6-dimensional formant feature reaches 93.5% and 94.1% accuracy on LJSpeech and TIMIT, while embedding configuration eigenvalues reach 88.1% and 87.2%.

Perspective

The results target binary voiced versus voiceless consonant classification, with experiments on public corpora including ALLSSTAR, LJSpeech, TIMIT, and LibriSpeech, using time-delay embedding plus persistent homology combined with traditional machine learning or a GRU decoder. For readers seeking interpretable features at lower computational cost, or improved classification stability under noise, TopCap and TopNN offer directly reusable pipelines; the cyclic time-delay embedding and parameter-selection analysis also apply to other time series requiring structural information.

Maximal persistence shows extreme sensitivity to the delay parameter while depending sublinearly (approximately as a square root) on embedding dimension, meaning parameter selection remains a key variable in practice; on real speech data, how points in the lower region of the diagram correspond to specific frequency, amplitude, or average-line changes is still unclear, and the authors note that vectorizing persistence diagrams to measure each fundamental variation remains a complex task; moreover, only maximal persistence and birth time are currently used, leaving the rest of the persistent homology information unexploited, with its potential gains yet to be explored.

Sources