A Study on Reducing Latency and Handling Accent Variation in Real-time Speech Translation Systems
DOI:
https://doi.org/10.54097/xe7av066Keywords:
Real-time Speech Translation, Latency Optimization, Accent Adaptation, Collaborative Optimization, Speech TechnologyAbstract
Real-time speech translation, or RST, sits at the heart of cross-language instant communication. It shows up in plenty of places now—cross-border video meetings, commercial translation gadgets, smart wearables, and so on. Still, two bottlenecks keep holding back practical deployment: latency control and accent robustness. Most existing reviews tend to zoom in on just one dimension, leaving the interplay between the two largely overlooked. This paper takes a systematic look at the core technologies and recent progress around latency optimization and accent adaptation in RST, covering work from 2020 to 2026. The approach draws on literature research, classification, and comparative analysis. Different technical paths get examined for what they do well and where they fall short. A recurring tension emerges—the trade-off between cutting latency and handling accent variation. The real difficulty lies in hitting a three-way balance: low latency, high accuracy, and strong accent robustness all at once. The paper also maps out shared research gaps in the field and points toward future directions worth pursuing. The aim is to offer some theoretical reference and practical guidance for the next wave of technical breakthroughs and real-world RST deployment.
Downloads
References
[1] Huawei Technologies Co., Ltd. (2023). A real time speech translation method based on lightweight CTC Transformer hybrid architecture (Chinese Patent No. ZL202310256789.0). China National Intellectual Property Administration.
[2] Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449 12460.
[3] Conneau, A., Lample, G., et al. (2022). Streaming transformer for low latency speech translation. Transactions of the Association for Computational Linguistics, 10, 890 905.
[4] Devlin, J., Chang, M. W., Lee, K., et al. (2023). Quantization aware training for transformer based speech translation models. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, 1245 1249.
[5] Zhou, S. Y., & Zhang, J. J. (2022). Adaptive chunk window for streaming speech translation with context awareness. Proceedings of the 2022 International Conference on Speech Communication and Processing, 789 793.
[6] Song, J., Shim, H., & Yang, E. (2021). Adaptively constrained monotonic multihead attention for streaming speech recognition. Proceedings of the 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 412 418. https://doi.org/10.1109/ASRU51503.2021.9688138. DOI: https://doi.org/10.1109/ASRU51503.2021.9688138
[7] Li, X., Wang, Y., & Zhang, H. (2026). Alignment based cascaded streaming translation for low latency cross language communication. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 34, 567 579.
[8] Yang, Z. L., Zhou, M., & Wang, H. F. (2023). Non standard accent data augmentation method based on speech synthesis. Journal of Chinese Information Processing, 37(4), 1 10.
[9] Castro, S., Jamshid Lou, P., et al. (2022). Domain adaptive pre training for multi accent speech recognition and translation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6789 6801.
[10] Lin, Y., Wang, H., & Li, X. (2024). Unsupervised cross lingual transfer learning for low resource accent speech recognition. Proceedings of the 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 8255–8259. https://doi.org/10.1109/ICASSP48485.2024.10446123. DOI: https://doi.org/10.1109/ICASSP48485.2024.10446123
[11] Liu, S., Chen, L., & Zhao, J. (2025). Few shot accent adaptation for real time speech translation using meta learning. Proceedings of the 2025 International Conference on Acoustics, Speech and Signal Processing, 2345 2349.
[12] ByteDance Inc. (2025). Seed LiveInterpret 2.0: An end to end low latency real time speech translation system. arXiv preprint arXiv:2501.08765.
[13] Chiu, C. C., & Raffel, C. (2018). Monotonic chunkwise attention for streaming neural networks. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 3744 3753.
[14] Rahmani, H., & Koudounas, A. (2024). Cross lingual transfer learning for low resource speech translation. HAL Open Science, hal 04432308.
[15] Wang, L., & Zhang, Q. (2024). StreamAtt: Direct streaming speech to text translation with attention based audio history selection. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 6123 6134.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Frontiers in Computing and Intelligent Systems

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

