☆ 3.8 Proceedings Paper

Chain-based Discriminative Autoencoders for Speech Recognition

INTERSPEECH 2022 (2022)

期刊

INTERSPEECH 2022

卷 -, 期 -, 页码 2078-2082

出版社

ISCA-INT SPEECH COMMUNICATION ASSOC

DOI: 10.21437/Interspeech.2022-10474

关键词

discriminative autoencoder; robust speech recognition; multi-condition training

类别

Acoustics Audiology & Speech-Language Pathology Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

In this paper, we propose three new versions of a discriminative autoencoder (DcAE) for speech recognition, achieving superior experimental results.

In our previous work, we proposed a discriminative autoencoder (DcAE) for speech recognition. DcAE combines two training schemes into one. First, since DcAE aims to learn encoder-decoder mappings, the squared error between the reconstructed speech and the input speech is minimized. Second, in the code layer, frame-based phonetic embeddings are obtained by minimizing the categorical cross-entropy between ground truth labels and predicted triphone-state scores. DcAE is developed based on the Kaldi toolkit by treating various TDNN models as encoders. In this paper, we further propose three new versions of DcAE. First, a new objective function that considers both categorical cross-entropy and mutual information between ground truth and predicted triphone-state sequences is used. The resulting DcAE is called a chain-based DcAE (c-DcAE). For application to robust speech recognition, we further extend c-DcAE to hierarchical and parallel structures, resulting in hc-DcAE and pc-DcAE. In these two models, both the error between the reconstructed noisy speech and the input noisy speech and the error between the enhanced speech and the reference clean speech are taken into the objective function. Experimental results on the WSJ and Aurora-4 corpora show that our DcAE models outperform baseline systems.

Chain-based Discriminative Autoencoders for Speech Recognition

期刊

INTERSPEECH 2022

出版社

ISCA-INT SPEECH COMMUNICATION ASSOC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Chain-based Discriminative Autoencoders for Speech Recognition

期刊

INTERSPEECH 2022

出版社

ISCA-INT SPEECH COMMUNICATION ASSOC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文