Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (8): 2533-2540.DOI: 10.11772/j.issn.1001-9081.2025070804
• Advanced computing • Previous Articles Next Articles
Jingying XUE1,2, Liqun KUANG1,2,3(
), Zhixun WANG1,2,3, Zhengtao GUO1,2,3, Huiyan HAN1,2,3
Received:2025-07-21
Revised:2025-09-30
Accepted:2025-10-14
Online:2025-11-05
Published:2026-08-10
Contact:
Liqun KUANG
About author:XUE Jingying, born in 2001, M. S. candidate. Her research interests include reinforcement learning, path planning.Supported by:
薛婧颖1,2, 况立群1,2,3(
), 王智巽1,2,3, 郭正涛1,2,3, 韩慧妍1,2,3
通讯作者:
况立群
作者简介:薛婧颖(2001—),女,山西运城人,硕士研究生,主要研究方向:强化学习、路径规划基金资助:CLC Number:
Jingying XUE, Liqun KUANG, Zhixun WANG, Zhengtao GUO, Huiyan HAN. Multi-agent path planning with hierarchical adaptive implicit quantiles[J]. Journal of Computer Applications, 2026, 46(8): 2533-2540.
薛婧颖, 况立群, 王智巽, 郭正涛, 韩慧妍. 分层自适应隐式分位数的多智能体路径规划[J]. 《计算机应用》唯一官方网站, 2026, 46(8): 2533-2540.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025070804
| 参数设置 | 值 | 参数设置 | 值 |
|---|---|---|---|
| 优化器 | Adam | 网络激活函数 | ReLU |
| 学习率 | 0.000 1 | 总时间步 | 6 000 000 |
| 折扣因子 | 0.99 | 批量大小 | 32 |
| 目标网络更新因子 | 1.0 | 探索率衰减步数 | 1 500 000 |
Tab. 1 Training hyperparameters
| 参数设置 | 值 | 参数设置 | 值 |
|---|---|---|---|
| 优化器 | Adam | 网络激活函数 | ReLU |
| 学习率 | 0.000 1 | 总时间步 | 6 000 000 |
| 折扣因子 | 0.99 | 批量大小 | 32 |
| 目标网络更新因子 | 1.0 | 探索率衰减步数 | 1 500 000 |
| 目标数 | 算法 | 智能体数为3 | 智能体数为4 | 智能体数为5 | 智能体数为6 | 智能体数为7 |
|---|---|---|---|---|---|---|
| 5 | Optimized DQN[ | 0.91 | 0.90 | 0.86 | 0.84 | 0.73 |
| HSP-DQN[ | 0.93 | 0.90 | 0.89 | 0.81 | 0.76 | |
| adaptive_IQN[ | 0.86 | 0.78 | ||||
| 本文算法 | 0.96 | 0.92 | 0.90 | |||
| 10 | Optimized DQN[ | 0.87 | 0.82 | 0.80 | 0.78 | 0.71 |
| HSP-DQN[ | 0.88 | 0.85 | 0.81 | 0.79 | 0.70 | |
| adaptive_IQN[ | 0.74 | |||||
| 本文算法 | 0.96 | 0.92 | 0.89 | 0.81 | ||
| 15 | Optimized DQN[ | 0.92 | 0.84 | 0.73 | 0.46 | 0.20 |
| HSP-DQN[ | 0.91 | 0.85 | 0.77 | 0.49 | 0.20 | |
| adaptive_IQN[ | ||||||
| 本文算法 | 0.97 | 0.91 | 0.81 | 0.57 | 0.27 | |
| 20 | Optimized DQN[ | 0.82 | 0.73 | 0.66 | 0.40 | 0.25 |
| HSP-DQN[ | 0.85 | 0.74 | 0.70 | 0.42 | 0.29 | |
| adaptive_IQN[ | ||||||
| 本文算法 | 0.90 | 0.81 | 0.73 | 0.49 | 0.33 |
Tab. 2 Comparison of success rate with different numbers of agents and targets
| 目标数 | 算法 | 智能体数为3 | 智能体数为4 | 智能体数为5 | 智能体数为6 | 智能体数为7 |
|---|---|---|---|---|---|---|
| 5 | Optimized DQN[ | 0.91 | 0.90 | 0.86 | 0.84 | 0.73 |
| HSP-DQN[ | 0.93 | 0.90 | 0.89 | 0.81 | 0.76 | |
| adaptive_IQN[ | 0.86 | 0.78 | ||||
| 本文算法 | 0.96 | 0.92 | 0.90 | |||
| 10 | Optimized DQN[ | 0.87 | 0.82 | 0.80 | 0.78 | 0.71 |
| HSP-DQN[ | 0.88 | 0.85 | 0.81 | 0.79 | 0.70 | |
| adaptive_IQN[ | 0.74 | |||||
| 本文算法 | 0.96 | 0.92 | 0.89 | 0.81 | ||
| 15 | Optimized DQN[ | 0.92 | 0.84 | 0.73 | 0.46 | 0.20 |
| HSP-DQN[ | 0.91 | 0.85 | 0.77 | 0.49 | 0.20 | |
| adaptive_IQN[ | ||||||
| 本文算法 | 0.97 | 0.91 | 0.81 | 0.57 | 0.27 | |
| 20 | Optimized DQN[ | 0.82 | 0.73 | 0.66 | 0.40 | 0.25 |
| HSP-DQN[ | 0.85 | 0.74 | 0.70 | 0.42 | 0.29 | |
| adaptive_IQN[ | ||||||
| 本文算法 | 0.90 | 0.81 | 0.73 | 0.49 | 0.33 |
| 目标数 | 智能体数 | Optimized DQN[ | HSP-DQN[ | adaptive_IQN[ | 本文方法 | ||||
|---|---|---|---|---|---|---|---|---|---|
| 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | ||
| 5 | 3 | 61.63 | 47.5 | 60.23 | 45.8 | 57.65 | 42.6 | ||
| 4 | 65.48 | 51.0 | 62.76 | 48.5 | 57.89 | 43.8 | |||
| 5 | 66.34 | 52.1 | 63.89 | 49.3 | 60.97 | 46.8 | |||
| 6 | 70.15 | 55.5 | 69.74 | 55.2 | 64.90 | 48.4 | |||
| 7 | 75.18 | 60.1 | 72.86 | 57.5 | 67.11 | 50.6 | |||
| 10 | 3 | 90.78 | 76.2 | 89.65 | 73.6 | 83.52 | 70.0 | ||
| 4 | 96.85 | 82.3 | 95.08 | 80.5 | 89.14 | 75.3 | |||
| 5 | 92.65 | 83.6 | 91.36 | 82.5 | 91.13 | 76.5 | |||
| 6 | 100.82 | 86.3 | 99.34 | 84.1 | 95.37 | 80.6 | |||
| 7 | 101.76 | 87.1 | 100.99 | 86.3 | 100.51 | 86.1 | |||
| 15 | 3 | 116.87 | 102.9 | 114.93 | 100.6 | 110.92 | 95.0 | ||
| 4 | 130.94 | 116.6 | 127.83 | 112.1 | 121.19 | 107.9 | |||
| 5 | 137.94 | 123.1 | 132.89 | 118.9 | 125.57 | 111.1 | |||
| 6 | 143.67 | 127.3 | 142.95 | 129.1 | 136.78 | 122.6 | |||
| 7 | 145.37 | 130.9 | 142.94 | 130.6 | 137.14 | 124.8 | |||
| 20 | 3 | 136.96 | 122.9 | 135.64 | 120.1 | 130.22 | 116.0 | ||
| 4 | 136.81 | 122.2 | 134.35 | 120.4 | 133.61 | 118.5 | |||
| 5 | 144.83 | 131.1 | 143.69 | 129.9 | 136.44 | 123.1 | |||
| 6 | 148.96 | 134.0 | 146.87 | 134.3 | 140.16 | 125.0 | |||
| 7 | 150.42 | 138.6 | 147.15 | 133.6 | 141.13 | 128.2 | |||
Tab. 3 Comparison of agent execution time and agent path length in different algorithms
| 目标数 | 智能体数 | Optimized DQN[ | HSP-DQN[ | adaptive_IQN[ | 本文方法 | ||||
|---|---|---|---|---|---|---|---|---|---|
| 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | 执行时间/s | 路径长度 | ||
| 5 | 3 | 61.63 | 47.5 | 60.23 | 45.8 | 57.65 | 42.6 | ||
| 4 | 65.48 | 51.0 | 62.76 | 48.5 | 57.89 | 43.8 | |||
| 5 | 66.34 | 52.1 | 63.89 | 49.3 | 60.97 | 46.8 | |||
| 6 | 70.15 | 55.5 | 69.74 | 55.2 | 64.90 | 48.4 | |||
| 7 | 75.18 | 60.1 | 72.86 | 57.5 | 67.11 | 50.6 | |||
| 10 | 3 | 90.78 | 76.2 | 89.65 | 73.6 | 83.52 | 70.0 | ||
| 4 | 96.85 | 82.3 | 95.08 | 80.5 | 89.14 | 75.3 | |||
| 5 | 92.65 | 83.6 | 91.36 | 82.5 | 91.13 | 76.5 | |||
| 6 | 100.82 | 86.3 | 99.34 | 84.1 | 95.37 | 80.6 | |||
| 7 | 101.76 | 87.1 | 100.99 | 86.3 | 100.51 | 86.1 | |||
| 15 | 3 | 116.87 | 102.9 | 114.93 | 100.6 | 110.92 | 95.0 | ||
| 4 | 130.94 | 116.6 | 127.83 | 112.1 | 121.19 | 107.9 | |||
| 5 | 137.94 | 123.1 | 132.89 | 118.9 | 125.57 | 111.1 | |||
| 6 | 143.67 | 127.3 | 142.95 | 129.1 | 136.78 | 122.6 | |||
| 7 | 145.37 | 130.9 | 142.94 | 130.6 | 137.14 | 124.8 | |||
| 20 | 3 | 136.96 | 122.9 | 135.64 | 120.1 | 130.22 | 116.0 | ||
| 4 | 136.81 | 122.2 | 134.35 | 120.4 | 133.61 | 118.5 | |||
| 5 | 144.83 | 131.1 | 143.69 | 129.9 | 136.44 | 123.1 | |||
| 6 | 148.96 | 134.0 | 146.87 | 134.3 | 140.16 | 125.0 | |||
| 7 | 150.42 | 138.6 | 147.15 | 133.6 | 141.13 | 128.2 | |||
| 指标 | Optimized DQN[ | HSP-DQN[ | adaptive_IQN[ | 本文算法 |
|---|---|---|---|---|
| 多目标动态适应性 | 0.86 | 0.89 | 0.94 | |
| 碰撞率 | 0.11 | 0.09 | 0.06 | |
| 路径冗余度 | 1.13 | 1.12 | 1.06 |
Tab. 4 Performance comparison in multi-objective dynamic adaptability and path conflict metrics
| 指标 | Optimized DQN[ | HSP-DQN[ | adaptive_IQN[ | 本文算法 |
|---|---|---|---|---|
| 多目标动态适应性 | 0.86 | 0.89 | 0.94 | |
| 碰撞率 | 0.11 | 0.09 | 0.06 | |
| 路径冗余度 | 1.13 | 1.12 | 1.06 |
| [1] | Zhang L, Cai Z, Yan Y, et al. Multi-agent policy learning-based path planning for autonomous mobile robots[J]. Engineering Applications of Artificial Intelligence, 2024, 129: No.107631. |
| [2] | Al-Kamil S J, Szabolcsi R. Optimizing path planning in mobile robot systems using motion capture technology[J]. Results in Engineering, 2024, 22: No.102043. |
| [3] | Yang S, Zhan X, Wang H. Low-communication collaborative navigation with state estimation for large-scale agent systems meeting minimum performance requirements[J]. Aerospace Science and Technology, 2025, 158: No.109940. |
| [4] | 李忠林,罗邵屏,贾玉婷. 移动机器人路径规划算法综述[J]. 现代信息科技, 2024, 8(19): 184-188. |
| Li Zhonglin, Luo Shaoping, Jia Yuting. Review of path planning algorithms for mobile robots[J]. Modern Information Technology, 2024, 8(19): 184-188. | |
| [5] | Ma H. Graph-based multi-robot path finding and planning[J]. Current Robotics Reports, 2022, 3(3): 77-84. |
| [6] | Bock S, Bomsdorf S, Boysen N, et al. A survey on the traveling salesman problem and its variants in a warehousing context[J]. European Journal of Operational Research, 2025, 322(1): 1-14. |
| [7] | Yang L, Li P, Qian S, et al. Path planning technique for mobile robots: a review[J]. Machines, 2023, 11(10): No.980. |
| [8] | Jin Y, Zhang Y, Yuan J, et al. Efficient multi-agent cooperative navigation in unknown environments with interlaced deep reinforcement learning[C]// ICASSP 2019. Piscataway: IEEE, 2019: 2897-2901. |
| [9] | Berseth G, Haworth B, Moon S, et al. Multi-agent hierarchical reinforcement learning for humanoid navigation[EB/OL]. (2023-05-06) [2025-04-29].. |
| [10] | Ndousse K, Eck D, Levine S, et al. Emergent social learning via multi-agent reinforcement learning[C]// Proceedings of the 38th International Conference on Machine Learning. New York: JMLR.org, 2021: 7991-8004. |
| [11] | Nekoei H, Badrinaaraayanan A, Sinha A, et al. Dealing with non-stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning[C]// Proceedings of the 2nd Conference on Lifelong Learning Agents. New York: JMLR.org, 2023: 376-398. |
| [12] | Zhao B, Jin W, Chen Z, et al. A semi-independent policies training method with shared representation for heterogeneous multi-agents reinforcement learning[J]. Frontiers in Neuroscience, 2023, 17: No.1201370. |
| [13] | Pina R, de Silva V, Artaud C, et al. Fully independent communication in multi-agent reinforcement learning[C]// AAMAS 2024. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, 2024: 2423-2425. |
| [14] | Park J Y, Lee J Y, Kim S B. Cooperative multi-agent reinforcement learning with approximate model learning[J]. IEEE Access, 2020, 8: 125389-125400. |
| [15] | Wang H, Wang X, Hu X, et al. A multi-agent reinforcement learning approach to dynamic service composition[J]. Information Sciences, 2016, 363: 96-119. |
| [16] | Khan A, Zhang C, Lee D D, et al. Scalable centralized deep multi-agent reinforcement learning via policy gradients[PP/OL]. V1. arXiv (2018-05-23) [2025-04-29].. |
| [17] | Chen S, Rui L, Gao Z, et al. Service migration with edge collaboration: multi-agent deep reinforcement learning approach combined with user preference adaptation[J]. Future Generation Computer Systems, 2025, 165: No.107612. |
| [18] | Yu C, Dong Y, Li Y, et al. Distributed multi-agent deep reinforcement learning for cooperative multi-robot pursuit[J]. The Journal of Engineering, 2020, 2020(13): 499-504. |
| [19] | Liu C, Tang F, Hu Y, et al. Distributed task migration optimization in MEC by extending multi-agent deep reinforcement learning approach[J]. IEEE Transactions on Parallel and Distributed Systems, 2021, 32(7): 1603-1614. |
| [20] | Guo S, Zhang X, Du Y, et al. Path planning of coastal ships based on optimized DQN reward function[J]. Journal of Marine Science and Engineering, 2021, 9(2): No.210. |
| [21] | Erkan E, Arserim M A. Mobile robot application with hierarchical start position DQN[J]. Computational Intelligence and Neuroscience, 2022, 2022: No.4115767. |
| [22] | Lin X, Huang Y, Chen F, et al. Decentralized multi-robot navigation for autonomous surface vehicles with distributional reinforcement learning[C]// ICRA 2024. Piscataway: IEEE, 2024: 8327-8333. |
| [1] | Juan CHEN, Yujie CHEN, Zongling WU, Di TIAN, Jie ZHONG. User-centric satellite edge computing architecture for task offloading optimization [J]. Journal of Computer Applications, 2026, 46(8): 2524-2532. |
| [2] | Yunle WANG, Xiang FENG, Huiqun YU. VLMDs-Privacy: privacy-enhanced strategy for cooperative decision-making in socially-aware multi-agent systems [J]. Journal of Computer Applications, 2026, 46(5): 1518-1525. |
| [3] | Shuai SU, Chenglin LIU. Observer-based leader-following consensus of heterogeneous descriptor multi-agent system with disturbance [J]. Journal of Computer Applications, 2026, 46(4): 1211-1217. |
| [4] | Caiqi WANG, Xining CUI, Yi XIONG, Shiqian WU. Adaptive extended RRT* path planning algorithm based on node-to-obstacle distance [J]. Journal of Computer Applications, 2025, 45(3): 920-927. |
| [5] | Xingwang WANG, Qingyang ZHANG, Shouyong JIANG, Yongquan DONG. Dynamic UAV path planning based on modified whale optimization algorithm [J]. Journal of Computer Applications, 2025, 45(3): 928-936. |
| [6] | Suqian WU, Jianguo YAN, Bin YANG, Tao QIN, Ying LIU, Jing YANG. Multi-strategy improved Aquila optimizer and its application in path planning [J]. Journal of Computer Applications, 2025, 45(3): 937-945. |
| [7] | Jing WANG, Xuming FANG. Intelligent joint power and channel allocation algorithm for Wi-Fi7 multi-link integrated communication and sensing [J]. Journal of Computer Applications, 2025, 45(2): 563-570. |
| [8] | Yu WANG, Mingyue ZHAO, Xiaolin ZHOU. Task-based assistive robot path planning in nursing home scenarios [J]. Journal of Computer Applications, 2025, 45(10): 3270-3276. |
| [9] | Yi RAN, Yongsheng LI, Ye JIANG. Addressing robot path planning issues using S-shaped growth curve integrated grasshopper optimization algorithm [J]. Journal of Computer Applications, 2025, 45(1): 178-185. |
| [10] | Baoyan SONG, Junxiang DING, Junlu WANG, Haolin ZHANG. Consortium blockchain modification method based on chameleon hash and verifiable secret sharing [J]. Journal of Computer Applications, 2024, 44(7): 2087-2092. |
| [11] | Tian MA, Runtao XI, Jiahao LYU, Yijie ZENG, Jiayi YANG, Jiehui ZHANG. Mobile robot 3D space path planning method based on deep reinforcement learning [J]. Journal of Computer Applications, 2024, 44(7): 2055-2064. |
| [12] | Runze TIAN, Yulong ZHOU, Hong ZHU, Gang XUE. Local information based path selection algorithm for service migration [J]. Journal of Computer Applications, 2024, 44(7): 2168-2174. |
| [13] | Xiaofang LIU, Jun ZHANG. Probability-driven dynamic multiobjective evolutionary optimization for multi-agent cooperative scheduling [J]. Journal of Computer Applications, 2024, 44(5): 1372-1377. |
| [14] | Jianqiang LI, Zhou HE. Hybrid NSGA-Ⅱ for vehicle routing problem with multi-trip pickup and delivery [J]. Journal of Computer Applications, 2024, 44(4): 1187-1194. |
| [15] | Zhaojun TANG, Meiyan XIA, Hua ZHANG, Ting XIE. Fixed-time consensus of dynamic event-triggered multi-agent systems [J]. Journal of Computer Applications, 2024, 44(3): 960-965. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||