用于语音增强的三阶段混合脉冲神经网络微调

Three-stage hybrid spiking neural networks fine-tuning for speech enhancement.

作者信息

Abuhajar Nidal, Wang Zhewei, Baltes Marc, Yue Ye, Xu Li, Karanth Avinash, Smith Charles D, Liu Jundong

机构信息

School of Electrical Engineering and Computer Science, Ohio University, Athens, OH, United States.

Department of Hearing, Speech, and Language Sciences, Ohio University, Athens, OH, United States.

出版信息

Front Neurosci. 2025 Apr 30;19:1567347. doi: 10.3389/fnins.2025.1567347. eCollection 2025.

DOI:10.3389/fnins.2025.1567347

PMID:40370668

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12075214/

Abstract

INTRODUCTION

In the past decade, artificial neural networks (ANNs) have revolutionized many AI-related fields, including Speech Enhancement (SE). However, achieving high performance with ANNs often requires substantial power and memory resources. Recently, spiking neural networks (SNNs) have emerged as a promising low-power alternative to ANNs, leveraging their inherent sparsity to enable efficient computation while maintaining performance.

METHOD

While SNNs offer improved energy efficiency, they are generally more challenging to train compared to ANNs. In this study, we propose a three-stage hybrid ANN-to-SNN fine-tuning scheme and apply it to Wave-U-Net and ConvTasNet, two major network solutions for speech enhancement. Our framework first trains the ANN models, followed by converting them into their corresponding spiking versions. The converted SNNs are subsequently fine-tuned with a hybrid training scheme, where the forward pass uses spiking signals and the backward pass uses ANN signals to enable backpropagation. In order to maintain the performance of the original ANN models, various modifications to the original network architectures have been made. Our SNN models operate entirely in the temporal domain, eliminating the need to convert wave signals into the spectral domain for input and back to the waveform for output. Moreover, our models uniquely utilize spiking neurons, setting them apart from many models that incorporate regular ANN neurons in their architectures.

RESULTS AND DISCUSSION

Experiments on noisy VCTK and TIMIT datasets demonstrate the effectiveness of the hybrid training, where the fine-tuned SNNs show significant improvement and robustness over the baseline models.

摘要

引言

在过去十年中，人工神经网络（ANN）彻底改变了许多与人工智能相关的领域，包括语音增强（SE）。然而，要使ANN实现高性能通常需要大量的功率和内存资源。最近，脉冲神经网络（SNN）作为一种有前途的低功耗替代方案出现，利用其固有的稀疏性在保持性能的同时实现高效计算。

方法

虽然SNN提供了更高的能源效率，但与ANN相比，它们通常更难训练。在本研究中，我们提出了一种三阶段的混合ANN到SNN微调方案，并将其应用于Wave-U-Net和ConvTasNet这两种语音增强的主要网络解决方案。我们的框架首先训练ANN模型，然后将它们转换为相应的脉冲版本。随后，使用混合训练方案对转换后的SNN进行微调，其中前向传播使用脉冲信号，反向传播使用ANN信号以实现反向传播。为了保持原始ANN模型的性能，对原始网络架构进行了各种修改。我们的SNN模型完全在时域中运行，无需将波形信号转换为频谱域进行输入，再转换回波形进行输出。此外，我们的模型独特地使用了脉冲神经元，这使它们与许多在架构中包含常规ANN神经元的模型有所不同。

结果与讨论

在有噪声的VCTK和TIMIT数据集上的实验证明了混合训练的有效性，其中经过微调的SNN相对于基线模型显示出显著的改进和鲁棒性。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/b5ff/12075214/b643c4d6c2ec/fnins-19-1567347-g0001.jpg

相似文献

Three-stage hybrid spiking neural networks fine-tuning for speech enhancement.用于语音增强的三阶段混合脉冲神经网络微调

Front Neurosci. 2025 Apr 30;19:1567347. doi: 10.3389/fnins.2025.1567347. eCollection 2025.

Spiking neural networks fine-tuning for brain image segmentation.用于脑图像分割的脉冲神经网络微调

Front Neurosci. 2023 Nov 1;17:1267639. doi: 10.3389/fnins.2023.1267639. eCollection 2023.

A universal ANN-to-SNN framework for achieving high accuracy and low latency deep Spiking Neural Networks.一种通用的 ANN-to-SNN 框架，可实现高精度和低延迟的深度尖峰神经网络。

Neural Netw. 2024 Jun;174:106244. doi: 10.1016/j.neunet.2024.106244. Epub 2024 Mar 15.

Enabling Spike-Based Backpropagation for Training Deep Neural Network Architectures.实现基于尖峰的反向传播以训练深度神经网络架构。

Front Neurosci. 2020 Feb 28;14:119. doi: 10.3389/fnins.2020.00119. eCollection 2020.

Rethinking the performance comparison between SNNS and ANNS.重新思考 SNNS 和 ANNS 的性能比较。

Neural Netw. 2020 Jan;121:294-307. doi: 10.1016/j.neunet.2019.09.005. Epub 2019 Sep 19.

Fast-SNN: Fast Spiking Neural Network by Converting Quantized ANN.快速脉冲神经网络：通过量化人工神经网络转换实现的快速脉冲神经网络

IEEE Trans Pattern Anal Mach Intell. 2023 Dec;45(12):14546-14562. doi: 10.1109/TPAMI.2023.3275769. Epub 2023 Nov 3.

Quantization Framework for Fast Spiking Neural Networks.快速脉冲神经网络的量化框架

Front Neurosci. 2022 Jul 19;16:918793. doi: 10.3389/fnins.2022.918793. eCollection 2022.

Efficient Processing of Spatio-Temporal Data Streams With Spiking Neural Networks.基于脉冲神经网络的时空数据流高效处理

Front Neurosci. 2020 May 5;14:439. doi: 10.3389/fnins.2020.00439. eCollection 2020.

Training much deeper spiking neural networks with a small number of time-steps.用少量时间步训练更深的尖峰神经网络。

Neural Netw. 2022 Sep;153:254-268. doi: 10.1016/j.neunet.2022.06.001. Epub 2022 Jun 15.

Backpropagation-Based Learning Techniques for Deep Spiking Neural Networks: A Survey.基于反向传播的深度学习尖峰神经网络学习技术综述。

IEEE Trans Neural Netw Learn Syst. 2024 Sep;35(9):11906-11921. doi: 10.1109/TNNLS.2023.3263008. Epub 2024 Sep 3.

本文引用的文献

Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip.基于尖峰的动态计算与异步传感计算神经形态芯片。

Nat Commun. 2024 May 25;15(1):4464. doi: 10.1038/s41467-024-47811-6.

Toward High-Accuracy and Low-Latency Spiking Neural Networks With Two-Stage Optimization.迈向具有两阶段优化的高精度低延迟脉冲神经网络

IEEE Trans Neural Netw Learn Syst. 2025 Feb;36(2):3189-3203. doi: 10.1109/TNNLS.2023.3337176. Epub 2025 Feb 6.

DIET-SNN: A Low-Latency Spiking Neural Network With Direct Input Encoding and Leakage and Threshold Optimization.DIET-SNN：一种具有直接输入编码以及泄漏和阈值优化的低延迟脉冲神经网络。

IEEE Trans Neural Netw Learn Syst. 2023 Jun;34(6):3174-3182. doi: 10.1109/TNNLS.2021.3111897. Epub 2023 Jun 1.

Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation.卷积时域音频分离网络（Conv-TasNet）：超越理想时频幅度掩蔽的语音分离方法

IEEE/ACM Trans Audio Speech Lang Process. 2019 Aug;27(8):1256-1266. doi: 10.1109/TASLP.2019.2915167. Epub 2019 May 6.

Going Deeper in Spiking Neural Networks: VGG and Residual Architectures.深入探索脉冲神经网络：VGG和残差架构。

Front Neurosci. 2019 Mar 7;13:95. doi: 10.3389/fnins.2019.00095. eCollection 2019.

Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks.用于训练高性能脉冲神经网络的时空反向传播

Front Neurosci. 2018 May 23;12:331. doi: 10.3389/fnins.2018.00331. eCollection 2018.

Conversion of Continuous-Valued Deep Networks to Efficient Event-Driven Networks for Image Classification.将连续值深度网络转换为用于图像分类的高效事件驱动网络

Front Neurosci. 2017 Dec 7;11:682. doi: 10.3389/fnins.2017.00682. eCollection 2017.

Training Deep Spiking Neural Networks Using Backpropagation.使用反向传播训练深度脉冲神经网络。

Front Neurosci. 2016 Nov 8;10:508. doi: 10.3389/fnins.2016.00508. eCollection 2016.

Unsupervised learning of digit recognition using spike-timing-dependent plasticity.使用基于脉冲时间依赖可塑性的无监督数字识别学习。

Front Comput Neurosci. 2015 Aug 3;9:99. doi: 10.3389/fncom.2015.00099. eCollection 2015.

Mapping from frame-driven to frame-free event-driven vision systems by low-rate rate coding and coincidence processing--application to feedforward ConvNets.通过低速率率编码和符合处理从基于帧的到无帧的事件驱动视觉系统的映射 - 应用于前馈 ConvNets。

IEEE Trans Pattern Anal Mach Intell. 2013 Nov;35(11):2706-19. doi: 10.1109/TPAMI.2013.71.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

用于语音增强的三阶段混合脉冲神经网络微调

Three-stage hybrid spiking neural networks fine-tuning for speech enhancement.

作者信息

机构信息

出版信息

INTRODUCTION

METHOD

RESULTS AND DISCUSSION

引言

方法

结果与讨论

相似文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

本文引用的文献