Research on a SAC-Based Algorithm for UAV Autonomous Navigation and Obstacle Avoidance
DOI:
https://doi.org/10.54097/yd2nna39Keywords:
Unmanned Aircraft, Obstacle Avoidance, SAC, Autonomous NavigationAbstract
The traditional reinforcement learning algorithms face the problems of poor complex spatial adaptability and inefficient exploration in the autonomous UAV pathfinding and obstacle avoidance tasks in three-dimensional unknown environments. To this end, this paper combines the dynamically decaying -greedy exploration strategy, the dynamic learning rate adjustment mechanism and the offline strategy SAC (Soft Actor-Critic) algorithm based on the maximum entropy reinforcement learning framework, and proposes an autonomous pathfinding and obstacle avoidance algorithm based on SAC, aiming to enhance the obstacle avoidance ability of UAVs in complex environments. -Greedy exploration strategy promotes exploration in the early stage of training and exploitation in the later stage, which helps the intelligent body to find the globally optimal strategy in complex environments and avoids falling into the local optimum in the early stage; the dynamic learning rate adjustment mechanism can refine the strategy and improve the training stability and final performance. The results in the test environment show that this algorithm makes the UAV collision rate decrease significantly, and the timeout rate is close to 0, which effectively enhances the algorithm's adaptability and stability to the environment.
Downloads
References
[1] Tao, Y., & Li, P. (2014). An overview of unmanned aircraft system development and key technologies. Aeronautical Manufacturing Technology,57(20),34-39.
[2] Liu, M., & Shi, H. (2025). Autonomous obstacle avoidance strategy for UAV based on EFRE-SAC. Computer System Applications,34(06),53-61.
[3] Wang, Z., & Xiang, X. (2018). Improved astar algorithm for path planning of marine robot. In 2018 37th chinese control conference (CCC), IEEE, 5410-5414.
[4] Prasad, N. L., & Ramkumar, B. (2022). 3-D deployment and trajectory planning for relay based UAV assisted cooperative communication for emergency scenarios using Dijkstra's algorithm. IEEE Transactions on Vehicular Technology, 72(4), 5049-5063.
[5] Du, Z., & Liu, S. (2018). Asymptotical RRT-based path planning for mobile robots in dynamic environments. In 2018 37th Chinese control conference (CCC), IEEE, 5281-5286.
[6] Chen, H., Chen, H., & Qiang, L. (2020). Multi-UAV 3D formation path planning based on improved artificial potential field. Journal of System Simulation, 32(3), 414-420.
[7] Sonny, A., Yeduri, S. R., & Cenkeramaddi, L. R. (2023). Q-learning-based unmanned aerial vehicle path planning with dynamic obstacle avoidance. Applied Soft Computing, 147, 110773.
[8] Qi, C., Wu, C., Lei, L., Li, X., & Cong, P. (2022). UAV path planning based on the improved PPO algorithm. In 2022 Asia Conference on Advanced Robotics, Automation, and Control Engineering (ARACE), IEEE, 193-199.
[9] Bouhamed, O., Ghazzai, H., Besbes, H., & Massoud, Y. (2020). Autonomous UAV navigation: A DDPG-based deep reinforcement learning approach. In 2020 IEEE International Symposium on circuits and systems (ISCAS), IEEE, 1-5.
[10] Luo, X., Wang, Q., Gong, H., & Tang, C. (2024). UAV path planning based on the average TD3 algorithm with prioritized experience replay. IEEE Access, 12, 38017-38029.
[11] Xue, B., Zhou, F., Wang, C., Gao, M., & Yin, L. (2024). Robot Mapless Navigation in VUCA Environments via Deep Reinforcement Learning. IEEE Transactions on Industrial Electronics, 72(1), 639-649.
[12] Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., ... & Levine, S. (2018). Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905.
[13] Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, Pmlr, 1861-1870.
[14] Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., ... & Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971.
[15] Fujimoto, S., Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods. In International conference on machine learning, PMLR, 1587-1596.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Frontiers in Computing and Intelligent Systems

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

