Tibetan Speech Emotion Recognition based on Capsule Network and Spatiotemporal Features
DOI:
https://doi.org/10.54097/84eqwt61Keywords:
Speech Emotion Recognition, Tibetan, Mel-frequency Cepstral Coefficients, Deep Learning, Capsule NetAbstract
Speech emotion recognition is an important branch of natural language processing that aims to automatically recognize and classify emotional information in speech through computer technology. In the specific language environment of Tibetan, due to relatively limited research and some existing studies appearing cumbersome and complex in feature extraction steps, a new network model has been proposed. The model is based on a capsule network and achieves lightweight design. It only uses Mel Frequency Cepstral Coefficients (MFCC) as its input features, extracts the spatiotemporal information of MFCC through multiple convolutional layers, and sends it into the capsule network for deep analysis. The recognition rate of 81.52% was achieved on the self-built Tibetan language emotion corpus TBSEC001. Meanwhile, the method achieved an unweighted accuracy (UA) of 85.63% and 95.54% respectively on the EMO-DB and RAVDESS public corpora, demonstrating the method's effectiveness.
Downloads
References
[1] Mencattini A, Martinelli E, Ringeval F, et al. Continuous estimation of emotions in speech by dynamic cooperative speaker models [J]. IEEE transactions on affective computing, 2016, 8 (3): 314-327.
[2] Hashem A, Arif M, Alghamdi M. Speech emotion recognition approaches: A systematic review [J]. Speech Communication, 2023: 102974.
[3] Al-Dujaili M J, Ebrahimi-Moghadam A. Speech emotion recognition: a comprehensive survey [J]. Wireless Personal Communications, 2023, 129 (4): 2525-2561.
[4] Atmaja B T, Akagi M. The effect of silence feature in dimensional speech emotion recognition [J]. arXiv preprint arXiv: 2003.01277, 2020.
[5] Akçay M B, Oğuz K. Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers [J]. Speech Communication, 2020, 116: 56-76.
[6] Fahad M S, Ranjan A, Yadav J, et al. A survey of speech emotion recognition in natural environment [J]. Digital signal processing, 2021, 110: 102951.
[7] Jain M, Narayan S, Balaji P, et al. Speech emotion recognition using support vector machine [J]. arXiv preprint arXiv: 2002. 07590, 2020.
[8] Aouani H, Ayed Y B. Speech emotion recognition with deep learning [J]. Procedia Computer Science, 2020, 176: 251-260.
[9] Wang J, Xue M, Culhane R, et al. Speech emotion recognition with dual-sequence LSTM architecture [C]//ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020: 6474-6478.
[10] Liu G, He W, Jin B. Feature fusion of speech emotion recognition based on deep learning [C]//2018 International conference on network infrastructure and digital content (IC-NIDC). IEEE, 2018: 193-197.
[11] Guo L, Wang L, Dang J, et al. A feature fusion method based on extreme learning machine for speech emotion recognition [C]//2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018: 2666 − 2670.
[12] Bandela S R, Kumar T K. Stressed speech emotion recognition using feature fusion of teager energy operator and MFCC [C]//2017 8th International Conference on Computing, Communication and Networking Technologies (ICCCNT). IEEE, 2017: 1-5.
[13] Lieskovská E, Jakubec M, Jarina R, et al. A review on speech emotion recognition using deep learning and attention mechanism [J]. Electronics, 2021, 10 (10): 1163.
[14] Niu Z, Zhong G, Yu H. A review on the attention mechanism of deep learning [J]. Neurocomputing, 2021, 452: 48-62.
[15] Vaswani A. Attention is all you need [J]. Advances in Neural Information Processing Systems, 2017.
[16] Ma J, Tang H, Zheng W L, et al. Emotion recognition using multimodal residual LSTM network [C]//Proceedings of the 27th ACM international conference on multimedia. 2019: 176-183.
[17] Xie Y, Liang R, Liang Z, et al. Speech emotion classification using attention-based LSTM [J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2019, 27 (11): 1675-1685.
[18] Pengmao Tashi, Cai Zhijie, Cai Rang Zhuoma. Construction of Tibetan emotional speech database [J]. Journal of Peking University (Natural Science Edition), 2023, 59 (05): 773-781. DOI: 10.13209/J.0479-8023.2022. 121.
[19] Gu Zeyue, Bianbawangdui, Qi Jindong. Tibetan speech emotion recognition based on multi-feature fusion [J]. Modern Electronic Technology, 2023, 46 (21): 129-133. DOI: 10. 16652/ J.issn.1004-373x. 2023.21. 024.
[20] Sabour S, Frosst N, Hinton G E. Dynamic routing between capsules [J]. Advances in neural information processing systems, 2017, 30.
[21] Lian Z, Liu B, Tao J. DECN: Dialogical emotion correction network for conversational emotion recognition [J]. Neurocomputing, 2021, 454: 483-495.
[22] Li S, Xing X, Fan W, et al. Spatiotemporal and frequential cascaded attention networks for speech emotion recognition [J]. Neurocomputing, 2021, 448: 238-248.
[23] Chen J, Liu Z. Mask dynamic routing to combined model of deep capsule network and u-net [J]. IEEE transactions on neural networks and learning systems, 2020, 31 (7): 2653-2664.
[24] Wen X C, Ye J X, Luo Y, et al. Ctl-mtnet: A novel capsnet and transfer learning-based mixed task net for the single-corpus and cross-corpus speech emotion recognition [J]. arXiv preprint arXiv: 2207.10644, 2022.
[25] Livingstone S R, Russo F A. The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English [J]. PloS one, 2018, 13 (5): e0196391.
[26] Qu Aitangs "Research on Tibetan Finals" [M] Qinghai Ethnic Publishing House. July 1991
[27] Yang Jie, Li Yonghong, Hu Axu, et al. Study on tone and laryngeal plug rhyme perception in Tibetan Lhasa dialect [J]. National Languages, 2023, (04): 101-110.
[28] Ye J, Wen X C, Wei Y, et al. Temporal modeling matters: A novel temporal emotional modeling approach for speech emotion recognition [C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.
[29] Sadok S, Leglaive S, Séguier R. A vector quantized masked autoencoder for speech emotion recognition [C]//2023 IEEE International conference on acoustics, speech, and signal processing workshops (ICASSPW). IEEE, 2023: 1-5.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Frontiers in Computing and Intelligent Systems

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

