A Deep Learning Based Framework for Blind Road Occupancy Detection

Authors

  • Yiteng Liu

DOI:

https://doi.org/10.54097/x7zrtx05

Keywords:

Machine learning, computer vision, blind road, PyBullet, convolutional neural network, vision transformer.

Abstract

This research paper addresses the pressing concern of road safety for the visually impaired in densely populated regions, especially in China. While tactile paving exists to guide blind individuals, unexpected obstacles pose serious hazards. The study proposes an Artificial Intelligence (AI) system utilizing advanced machine learning techniques to identify obstacles on blind roads, providing real-time feedback for navigation. To overcome data limitations, a virtual environment using Pybullet is created for data generation, combining synthetic and real-world images for training. This study introduces a Convolutional Neural Network (CNN)-based model and integrates a Vision Transformer (ViT) model, comparing their efficacies. The progressive training approach yields a highly effective CNN model, outperforming ViT models. The practical application of the CNN model in real-world scenarios has proven its efficacy in detecting obstacles, underscoring its reliability and significant contribution to improving road safety for visually impaired individuals. This research extends beyond providing a safety enhancement measure; it also sheds light on the wider applicability of such methods in addressing issues like the scarcity of real-world data.

Downloads

Download data is not yet available.

References

Ma X, Dai Z, He Z, et al. Learning traffic as images: A deep convolutional neural network for large-scale transportation network speed prediction. Sensors, 2017, 17(4): 818.

Redmon J, Farhadi A. YOLO9000: better, faster, stronger. Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 7263-7271.

Kavraki L E, Svestka P, Latombe J C, et al. Probabilistic roadmaps for path planning in high-dimensional configuration spaces. IEEE transactions on Robotics and Automation, 1996, 12(4): 566-580.

Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 2015, 28.

He K, Gkioxari G, Dollár P, et al. Mask r-cnn. Proceedings of the IEEE international conference on computer vision. 2017: 2961-2969.

Pybullet quickstart guide [EB/OL]. Google Docs, Google, 2023, https://docs.google.com/document/d/10sXEhzFRSnvFcl3XxNGhnD4N2SedqwdAvK3dsihxVUA/edit#heading=h.2ye70wns7io3.

Sgss8. Standard Size of Tactile Paving Bricks, Dimensions of Tactile Paving Bricks [EB/OL]. Sad Sayings Bar, 2023, https://www.sgss8.net/tpdq/7454697/1.htm.

Fukushima K. Neocognitron: A hierarchical neural network capable of visual pattern recognition. Neural networks, 1988, 1(2): 119-130.

LeCun Y. Generalization and network design strategies. Connectionism in perspective, 1989, 19(143-155): 18.

LeCun Y, Boser B, Denker J S, et al. Backpropagation applied to handwritten zip code recognition. Neural computation, 1989, 1(4): 541-551.

LeCun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998, 86(11): 2278-2324.

Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 2012, 25.

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in neural information processing systems, 2017, 30.

Liu L, Liu X, Gao J, et al. Understanding the difficulty of training transformers. arXiv preprint arXiv:2004.08249, 2020.

Downloads

Published

26-04-2024

How to Cite

Liu, Y. (2024). A Deep Learning Based Framework for Blind Road Occupancy Detection. Highlights in Science, Engineering and Technology, 94, 355-367. https://doi.org/10.54097/x7zrtx05