☆ 3.8 Proceedings Paper

LEARNING TO ENHANCE OR NOT: NEURAL NETWORK-BASED SWITCHING OF ENHANCED AND OBSERVED SIGNALS FOR OVERLAPPING SPEECH RECOGNITION

2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP) (2022)

期刊

2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP)

卷 -, 期 -, 页码 6287-6291

出版社

IEEE

DOI: 10.1109/ICASSP43922.2022.9746347

关键词

input switching; speech extraction; speech separation; noise-robust speech recognition; SpeakerBeam

类别

Acoustics Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

In this work, a DNN-based switching method is proposed to address the performance degradation issue caused by speech enhancement in automatic speech recognition (ASR). By learning to estimate the suitability of enhanced or observed signals for ASR, and using soft-switching to combine them, the proposed method achieves improved ASR performance.

The combination of a deep neural network (DNN) -based speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end is a widely used approach to implement overlapping speech recognition. However, the SE front-end generates processing artifacts that can degrade the ASR performance. We previously found that such performance degradation can occur even under fully overlapping conditions, depending on the signal-to-interference ratio (SIR) and signal-to-noise ratio (SNR). To mitigate the degradation, we introduced a rule-based method to switch the ASR input between the enhanced and observed signals, which showed promising results. However, the rule's optimality was unclear because it was heuristically designed and based only on SIR and SNR values. In this work, we propose a DNN-based switching method that directly estimates whether ASR will perform better on the enhanced or observed signals. We also introduce soft-switching that computes a weighted sum of the enhanced and observed signals for ASR input, with weights given by the switching model's output posteriors. The proposed learning-based switching showed performance comparable to that of rule-based oracle switching. The soft-switching further improved the ASR performance and achieved a relative character error rate reduction of up to 23 % as compared with the conventional method.

LEARNING TO ENHANCE OR NOT: NEURAL NETWORK-BASED SWITCHING OF ENHANCED AND OBSERVED SIGNALS FOR OVERLAPPING SPEECH RECOGNITION

期刊

2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP)

出版社

IEEE

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

LEARNING TO ENHANCE OR NOT: NEURAL NETWORK-BASED SWITCHING OF ENHANCED AND OBSERVED SIGNALS FOR OVERLAPPING SPEECH RECOGNITION

期刊

2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP)

出版社

IEEE

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文