A Review of Street Scene Semantic Segmentation: Current Status, Challenges, and Future Prospects

Authors

  • Linlin Liu

DOI:

https://doi.org/10.54097/h97y2f72

Keywords:

Deep learning; Semantic segmentation; Autonomous driving

Abstract

Street scene semantic segmentation is a core task in environmental perception for autonomous driving, which has made significant progress in recent years driven by deep learning. This paper systematically reviews the current research status in this field, analyzes technical challenges, and explores future development directions. First, it examines the technological evolution of traditional machine learning and deep learning methods, focusing on model innovations based on Convolutional Neural Networks (CNNs) and Transformer architectures. Then, it discusses cutting-edge solutions such as multimodal fusion, self-supervised learning, and 3D segmentation to address key challenges, including complex scene understanding, real-time constraints, and data efficiency. Finally, in the context of smart transportation and urban digital management, it proposes pathways to overcome technical bottlenecks, providing theoretical references for future research and practical applications.

References

[1] Feng D, Haase-Schütz C, Rosenbaum L, et al. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges[J]. IEEE Transactions on Intelligent Transportation Systems, 2020, 22(3): 1341-1360.

[2] Treml M, Arjona-Medina J, Unterthiner T, et al. Speeding up semantic segmentation for autonomous driving[J]. 2016.

[3] Garcia-Garcia A, Orts-Escolano S, Oprea S, et al. A review on deep learning techniques applied to semantic segmentation[J]. arXiv preprint arXiv:1704.06857, 2017.

[4] Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks[J]. Advances in neural information processing systems, 2012, 25.

[5] Jiang F, Grigorev A, Rho S, et al. Medical image semantic segmentation based on deep learning[J]. Neural Computing and Applications, 2018, 29: 1257-1265.

[6] Pohle R, Toennies K D. Segmentation of medical images using adaptive region growing[C]//Medical Imaging 2001: Image Processing. SPIE, 2001, 4322: 1337-1346.

[7] Halder A, Dey D, Sadhu A K. Lung nodule detection from feature engineering to deep learning in thoracic CT images: a comprehensive review[J]. Journal of digital imaging, 2020, 33(3): 655-677.

[8] Seghers D, Hermans J, Loeckx D, et al. Model-based segmentation using graph representations[C]//Medical Image Computing and Computer-Assisted Intervention–MICCAI 2008: 11th International Conference, New York, NY, USA, September 6-10, 2008, Proceedings, Part I 11. Springer Berlin Heidelberg, 2008: 393-400.

[9] Chen S, Gamechi Z S, Dubost F, et al. An end-to-end approach to segmentation in medical images with CNN and posterior-CRF[J]. Medical image analysis, 2022, 76: 102311.

[10] Dai J, Li Y, He K, et al. R-fcn: Object detection via region-based fully convolutional networks[J]. Advances in neural information processing systems, 2016, 29.

[11] Sahli H, Ben Slama A, Labidi S. U-Net: A valuable encoder-decoder architecture for liver tumors segmentation in CT images[J]. Journal of X-ray science and technology, 2022, 30(1): 45-56.

[12] Chen L C, Papandreou G, Schroff F, et al. Rethinking atrous convolution for semantic image segmentation[J]. arXiv preprint arXiv:1706.05587, 2017.

[13] Chen L C, Papandreou G, Kokkinos I, et al. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 40(4): 834-848.

[14] Liu T, Luo R, Xu L, et al. Spatial channel attention for deep convolutional neural networks[J]. Mathematics, 2022, 10(10): 1750.

[15] Feng H, Liu W, Xu H, et al. A lightweight dual-branch semantic segmentation network for enhanced obstacle detection in ship navigation[J]. Engineering Applications of Artificial Intelligence, 2024, 136: 108982.

[16] Wang W, Dai J, Chen Z, et al. Internimage: Exploring large-scale vision foundation models with deformable convolutions[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 14408-14419.

[17] Yuan Y, Fu R, Huang L, et al. Hrformer: High-resolution vision transformer for dense predict[J]. Advances in neural information processing systems, 2021, 34: 7281-7293.

Downloads

Published

26-03-2025

Issue

Section

Articles

How to Cite

Liu, L. (2025). A Review of Street Scene Semantic Segmentation: Current Status, Challenges, and Future Prospects. Mathematical Modeling and Algorithm Application, 4(2), 33-38. https://doi.org/10.54097/h97y2f72