拓扑增强的语音信号处理机器学习方法
核心概要
该工作提出 TopCap 与 TopNN 两种方法,用时间延迟嵌入与持续同调从语音时间序列中提取拓扑特征(如最大持续度及其出生时间),在浊辅音与清辅音二分类任务上,TopCap 在小型数据集上取得与部分先进神经网络相当的准确率并具备更高效率与可解释性,而将拓扑特征与门控循环单元拼接的 TopNN 在多个数据集及不同信噪比噪声条件下比标准神经网络取得更高准确率、更稳定表现与更强抗噪能力。
Figure 1: Visualisation for data and some of their commonly used topological descriptors. a Time-delay embedding (dimension=3, delay=10, skip=1) of f ( t n ) = sin ( 2 t n ) − 3 sin ( t n ) f(t_{n})=\sin(2t_{n})-3\sin(t_{n}) , with t n = π 50 n t_{n}=\frac{\pi}{50}n ( 0 ≤ n ≤ 200 0\leq n\leq 200 ). Resulting point clouds lay on a closed curve in 3-dimensional Euclidean space. The colour indicates original locations of data in the time series. b A topological space and its triangulation. On the left is a topological space consisting of a 1-dimensional sphere (i.e., a circle) and a 2-dimensional sphere with a single point of contact, denoted as 𝕊 1 ∨ 𝕊 2 \mathbb{S}^{1}\vee\mathbb{S}^{2} . The right depicts a triangulation of this topological space. c Average temperature in the U.S. with monthly values (dark blue dots) and yearly values (green curve). The left panel shows a single-year section of average temperature. d Computing PH. The four plots consecutively show how a persistence diagram or barcode is computed: Connect each pair of points with a distance less than ϵ \epsilon by a line segment, fill in each triple of points with mutual distances less than ϵ \epsilon with a triangular region, etc., and compute the corresponding homology groups. e Characterising the vibration of a time series in terms of its variability of frequency, amplitude, and average line. f Commonly used representations for PH, with an example of 100 points uniformly distributed over a bounded region in 2D Euclidean space. A persistence barcode is a multiset of intervals, where the horizontal axis shows when each feature appears and disappears. A persistence diagram directly plots the birth and death times (values of ϵ \epsilon , not to be confused with time series) of each interval. In both plots, 0 and 1 correspond to the 0-dimensional loops (connected components) and 1-dimensional loops. In a persistence landscape, the k k th landscape is the k k th largest value of tent functions for each feature, here taken as the 1-dimensional loops, with the horizontal axis representing resolution (turning the persistence diagram clockwise by 45 ∘ 45^{\circ} ). Similarly, a persistence image is created by applying Gaussian functions centred at each feature and then converting them into a pixelated image, where both the horizontal and vertical axes represent resolution.
arXiv深度剖析
提出 TopCap:以时间延迟嵌入(维度 d=100、延迟 τ=6T/d)结合持续同调,从一维持续图中提取最大持续度及其出生时间作为特征,输入树、判别、回归、朴素贝叶斯、支持向量机、k 近邻、线性与集成等传统分类器。 相较以往依赖能量与频谱信息(STFT、MFCC)的语音处理方法,该流程直接刻画时间序列的拓扑结构,且仅用最大持续度与出生时间两个量即可区分浊辅音与清辅音。 在 5101 条记录(3571 训练、1530 测试,其中浊辅音 2138 条、清辅音 2963 条)上,多数算法 AUC 超过 96%、准确率超过 93%;浊辅音倾向具有更高的出生时间与持续度。
在 8 个小型数据集与 4 个大型数据集上,将 TopCap 与 STFT–CNN、MFCC–GRU、MFCC–Transformer 对比。 TopCap 在小型数据集上超过 MFCC–GRU 与 STFT–CNN-8,在大型数据集上总体不及深度神经网络,这一差距促成了拓扑特征与神经网络结合的设计。 小型数据集准确率如 HT1 上 94.3%、LJ 上 94.6%、TIMIT 上 83.9%;大型数据集如 ALLSSTAR 上 92.5%、LJSpeech 上 92.9%、TIMIT 上 92.8%、LibriSpeech 上 88.7%。
提出 TopNN:将 TopCap 的最大持续度特征与 GRU 的 6 维隐状态拼接为 7 维向量,经全连接解码器分类,并以 ZeroNN(拓扑特征置零)作为对照。 该设计使性能差异可归因于拓扑特征的加入,而非模型容量变化;在原始与加噪数据上 TopNN 均优于 ZeroNN 与 NN,且噪声越强差距越明显。 在 ALLSSTAR、LJSpeech、TIMIT 上,TopNN 在无噪声至强噪声(SNR=0dB)各档均取得更高训练与测试准确率,且多次实验的准确率方差更低。
在合成与真实数据上探索持续图对非周期性振动模式的刻画能力,并提出共振峰频谱特征与循环时间延迟嵌入配置特征值作为额外几何特征。 合成实验显示持续图低持续区域的点分布可区分频率、振幅与平均线三类基本变化;真实语音中不稳定序列在持续图低区点密度更高,稳定序列倾向获得高最大持续度。 合成数据使用维度 100、延迟 3、跳步 10 的参数设置;真实数据取自 ALLSSTAR 中两个元音 [A] 录音,按 600、800、1000、1200 四个重叠区间截取;共振峰 6 维特征在 LJSpeech 与 TIMIT 上分别取得 93.5% 与 94.1% 准确率,嵌入配置特征值分别取得 88.1% 与 87.2%。
启示与展望
该结果面向浊辅音与清辅音的二分类任务,实验数据来自 ALLSSTAR、LJSpeech、TIMIT 与 LibriSpeech 等公开语料,方法设定为时间延迟嵌入加持续同调并配合传统机器学习或 GRU 解码器。对希望以较低计算成本获得可解释特征、或在噪声条件下提升分类稳定性的读者,TopCap 与 TopNN 提供了可直接参照的流程;循环时间延迟嵌入与参数选择分析也可用于其他需要结构信息的时间序列。
最大持续度对延迟参数表现出极端敏感性,而对嵌入维度呈近似平方根的次线性依赖,这意味着参数选择仍是实际使用中的关键变量;在真实语音数据上,持续图低区点的分布如何对应到具体的频率、振幅或平均线变化尚不清楚,作者也指出向量化持续图以度量各类基本变化仍是复杂任务;此外,拓扑特征目前仅取最大持续度与出生时间两个量,其余持续同调信息未被利用,其潜在增益有待进一步探索。
