Research on Data Augmentation and Deep Learning Algorithms for Speech Emotion Computing

Authors

  • Kaiyu Guo Xizang Agricultural and Animal Husbandry University, Linzhi, China; Xizang Guoang Information Technology Co., Ltd., Lasa, China
  • Yungang Wei Xizang University, Lasa, China

DOI:

https://doi.org/10.54097/pc8s9673

Keywords:

Speech Emotion Computing, Data Augmentation, Deep Learning, Computational Linguistics, Feature Extraction, Emotion Recognition

Abstract

Speech is the most natural human interaction medium, which contains both linguistic semantic information and numerous non-semantic linguistic characteristics such as emotion, attitude and mood, making it a core research object in the fields of computational linguistics and human-computer interaction. Speech emotion computing aims to explore emotional features embedded in speech signals through computer algorithms to realize automatic recognition and classification of human emotions. It is an interdisciplinary research area integrating linguistics, signal processing and artificial intelligence. At present, mainstream research on speech emotion recognition generally faces problems including insufficient annotated corpora, single sample dimensions, poor model generalization ability and low recognition accuracy under low-resource conditions. This paper systematically conducts research from two major perspectives: data augmentation and deep learning algorithms. Firstly, it sorts out the computational linguistic theoretical basis of speech emotion computing and analyzes the linguistic implications of prosodic, spectral and temporal emotional features in speech signals. Secondly, it investigates traditional data augmentation and intelligent deep learning-based augmentation methods separately, and optimizes augmentation strategies according to the linguistic characteristics of oral speech corpora. In addition, a deep learning recognition model suitable for speech emotional features is constructed to improve the feature extraction capability and classification performance of the model. Finally, extensive comparative experiments are carried out to verify the joint optimization effect of the proposed data augmentation algorithms and deep learning models. Experimental results demonstrate that the improved multi-dimensional data augmentation method can effectively enrich speech corpora and address the shortage of low-resource annotated data. Combined with the optimized deep learning model, the proposed method significantly improves the accuracy and stability of speech emotion recognition, providing technical support for natural language interaction scenarios such as intelligent customer service, human-computer dialogue and emotion monitoring.

Downloads

Download data is not yet available.

References

[1] Kragel, P. A., Reddan, M. C., Labar, K. S., et al. (2019). Emotion schemas are embedded in the human visual system. Science Advances, 5(7), 1–15.

[2] Ankit, K., Yu, Z. W., Hillary, S., et al. (2021). A preliminary exploration of virtual reality-based visual and touch sensory processing assessment for adolescents with autism spectrum disorder. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 29(1), 619–628.

[3] Liu, T. J., Li, F. L., & Jiang, Y. (2017). Cortical dynamic causality network for auditory-motor tasks. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 25(8), 1092–1099.

[4] Davis, K., Biddulph, R., & Balashek, S. (1952). Automatic recognition of spoken digits. Journal of the Acoustical Society of America, 24(6), 637–644.

[5] Jelinek, F., Bahl, L., & Mercer, R. (1975). Design of a linguistic statistical decoder for the recognition of continuous speech. IEEE Transactions on Information Theory, 21(3), 250–256.

Downloads

Published

16 August 2026

Issue

Section

Articles

How to Cite

Guo, K., & Wei, Y. (2026). Research on Data Augmentation and Deep Learning Algorithms for Speech Emotion Computing. International Journal of Education and Humanities, 24(2), 91-94. https://doi.org/10.54097/pc8s9673