Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2725-2731.DOI: 10.11772/j.issn.1001-9081.2025081006
• Artificial intelligence •
Tao XU1,2, Bin HU1,2, Jin QIN1,2(
)
Received:2025-09-04
Revised:2026-01-08
Accepted:2026-02-03
Online:2026-02-12
Published:2026-09-10
Contact:
Jin QIN
About author:XU Tao, born in 2000, M. S. candidate. His research interests include reinforcement learning.Supported by:通讯作者:
秦进
作者简介:许涛(2000—),男,贵州遵义人,硕士研究生,主要研究方向:强化学习基金资助:CLC Number:
Tao XU, Bin HU, Jin QIN. Maximum entropy reinforcement learning method with temperature coefficient adaptive adjustment[J]. Journal of Computer Applications, 2026, 46(9): 2725-2731.
许涛, 胡滨, 秦进. 温度系数自适应调节的最大熵强化学习方法[J]. 《计算机应用》唯一官方网站, 2026, 46(9): 2725-2731.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025081006
| 控制任务 | 动作维度 | 状态维度 | 任务目标 |
|---|---|---|---|
| HalfCheetah-v2 | 6 | 17 | 控制猎豹机器人前进 |
| Ant-v2 | 8 | 111 | 控制四足机器人前爬 |
| Hopper-v2 | 3 | 11 | 控制单腿前跳 |
| Walker2d-v2 | 6 | 17 | 控制双腿步态平衡 |
| Humanoid-v2 | 17 | 376 | 控制人形机器人行走 |
Tab. 1 Action/state dimensions and task objectives for different tasks
| 控制任务 | 动作维度 | 状态维度 | 任务目标 |
|---|---|---|---|
| HalfCheetah-v2 | 6 | 17 | 控制猎豹机器人前进 |
| Ant-v2 | 8 | 111 | 控制四足机器人前爬 |
| Hopper-v2 | 3 | 11 | 控制单腿前跳 |
| Walker2d-v2 | 6 | 17 | 控制双腿步态平衡 |
| Humanoid-v2 | 17 | 376 | 控制人形机器人行走 |
| 训练参数 | 值 | 训练参数 | 值 |
|---|---|---|---|
| Actor网络学习率 | 0.000 3 | 固定温度系数 | 0.2 |
| Critic网络学习率 | 0.000 3 | 更新间隔 | 1 |
| 折扣率 | 0.99 | 噪声标准差 | 0.2 |
| 训练批次大小 | 256 | 测试轮数 | 10 |
| 经验池大小 | 1 000 000 | 裁剪范围-TD3 | 0.5 |
| 软更新参数 | 0.005 | 裁剪范围-PPO | 0.2 |
Tab. 2 Main training parameters of different algorithms
| 训练参数 | 值 | 训练参数 | 值 |
|---|---|---|---|
| Actor网络学习率 | 0.000 3 | 固定温度系数 | 0.2 |
| Critic网络学习率 | 0.000 3 | 更新间隔 | 1 |
| 折扣率 | 0.99 | 噪声标准差 | 0.2 |
| 训练批次大小 | 256 | 测试轮数 | 10 |
| 经验池大小 | 1 000 000 | 裁剪范围-TD3 | 0.5 |
| 软更新参数 | 0.005 | 裁剪范围-PPO | 0.2 |
| 任务 | PPO | DDPG | TD3 | SAC-A | SAC-F | STAA-SAC |
|---|---|---|---|---|---|---|
| HalfCheetah-v2 | 2 287±57 | 9 399±185 | 10 201±75 | 10 432±122 | 11 602±99 | 10 887±62 |
| Ant-v2 | 1 944±88 | 1 300±190 | 4 368±89 | 4 776±143 | 5 260±170 | 5 723±84 |
| Hopper-v2 | 2 702±214 | 2 186±201 | 2 917±241 | 3 145±152 | 3 357±31 | 3 459±10 |
| Walker2d-v2 | 2 408±127 | 2 034±254 | 3 528±326 | 4 528±355 | 4 007±216 | 4 722±108 |
| Humanoid-v2 | 447±13 | 178±18 | 4 938±120 | 4 733±391 | 4 834±378 | 5 242±205 |
Tab. 3 Performance mean and standard deviation of final 10% data segments in MuJoCo test environment
| 任务 | PPO | DDPG | TD3 | SAC-A | SAC-F | STAA-SAC |
|---|---|---|---|---|---|---|
| HalfCheetah-v2 | 2 287±57 | 9 399±185 | 10 201±75 | 10 432±122 | 11 602±99 | 10 887±62 |
| Ant-v2 | 1 944±88 | 1 300±190 | 4 368±89 | 4 776±143 | 5 260±170 | 5 723±84 |
| Hopper-v2 | 2 702±214 | 2 186±201 | 2 917±241 | 3 145±152 | 3 357±31 | 3 459±10 |
| Walker2d-v2 | 2 408±127 | 2 034±254 | 3 528±326 | 4 528±355 | 4 007±216 | 4 722±108 |
| Humanoid-v2 | 447±13 | 178±18 | 4 938±120 | 4 733±391 | 4 834±378 | 5 242±205 |
| [1] | Wang J, Karatzoglou A, Arapakis I, et al. Reinforcement learning-based recommender systems with large language models for state reward and action modeling [C]// SIGIR 2024. New York: ACM, 2024: 375-385. |
| [2] | 代珊珊,刘全. 基于动作约束深度强化学习的安全自动驾驶方法[J]. 计算机科学, 2021, 48(9): 235-243. |
| Dai Shanshan, Liu Quan. Action constrained deep reinforcement learning based safe automatic driving method [J]. Computer Science, 2021, 48(9): 235-243. | |
| [3] | Wang X, Zhang J, Hou D, et al. Autonomous driving based on approximate safe action [J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(12): 14320-14328. |
| [4] | 何浩东,符浩,王强,等. 基于深度强化学习的多机器人路径跟随与编队[J]. 计算机应用, 2024, 44(8): 2626-2633. |
| He Haodong, Fu Hao, Wang Qiang, et al. Multi-robot path following and formation based on deep reinforcement learning[J]. Journal of Computer Applications, 2024, 44(8): 2626-2633. | |
| [5] | 马天,席润韬,吕佳豪,等. 基于深度强化学习的移动机器人三维路径规划方法[J]. 计算机应用, 2024, 44(7): 2055-2064. |
| Ma Tian, Xi Runtao, Jiahao Lyu, et al. Mobile robot 3D space path planning method based on deep reinforcement learning [J]. Journal of Computer Applications, 2024, 44(7): 2055-2064. | |
| [6] | Haarnoja T, Tang H, Abbeel P, et al. Reinforcement learning with deep energy-based policies [J]. Proceedings of Machine Learning Research, 2017, 70: 1352-1361. |
| [7] | Ziebart B D, Maas A, Bagnell J A, et al. Maximum entropy inverse reinforcement learning [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2008, 3: 1433-1438. |
| [8] | Boularias A, Kober J, Peters J. Relative entropy inverse reinforcement learning [J]. Proceedings of Machine Learning Research, 2011, 15: 182-189. |
| [9] | Ng A Y, Russell S J. Algorithms for inverse reinforcement learning[C]// ICML 2000. San Francisco: Morgan Kaufmann Publishers Inc., 2000: 663-670. |
| [10] | Nachum O, Norouzi M, Xu K, et al. Bridging the gap between value and policy based reinforcement learning [C]// NeurIPS 2017. Red Hook: Curran Associates Inc., 2017: 2772-2782. |
| [11] | Mnih V, Badia A P, Mirza M, et al. Asynchronous methods for deep reinforcement learning [J]. Proceedings of Machine Learning Research, 2016, 48: 1928-1937. |
| [12] | Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor [J]. Proceedings of Machine Learning Research, 2018, 80: 1861-1870. |
| [13] | Lillicrap T P, Hunt J J, Pritzel A, et al. Continuous control with deep reinforcement learning[PP/OL]. V6. arXiv (2019-07-05) [2025-08-13].. |
| [14] | Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms[PP/OL]. V2. arXiv (2017-08-28) [2025-08-26].. |
| [15] | Fujimoto S, van Hoof H, Meger D. Addressing function approximation error in actor-critic methods [J]. Proceedings of Machine Learning Research, 2018, 80: 1587-1596. |
| [16] | Todorov E, Erez T, Tassa Y. MuJoCo: a physics engine for model-based control [C]// IROS 2012. Piscataway: IEEE, 2012: 5026-5033. |
| [17] | Haarnoja T, Zhou A, Hartikainen K, et al. Soft actor-critic algorithms and applications[PP/OL]. V2. arXiv (2019-01-29) [2025-08-17].. |
| [18] | Igoe C, Pande S, Venkatraman S, et al. Multi-alpha soft actor-critic: overcoming stochastic biases in maximum entropy reinforcement learning [C]// ICRA 2023. Piscataway: IEEE, 2023: 7162-7168. |
| [19] | Wang Y, Ni T. Meta-SAC: auto-tune the entropy temperature of soft actor-critic via metagradient[PP/OL]. V2. arXiv (2020-07-31) [2025-07-17].. |
| [20] | Sutton R S. Learning to predict by the methods of temporal differences [J]. Machine Learning, 1988, 3(1): 9-44. |
| [21] | Ding Z, Huang Y, Yuan H, et al. Introduction to reinforcement learning [M]// Dong H, Ding Z, Zhang S. Deep reinforcement learning: fundamentals, research and applications. Singapore: Springer, 2020: 47-123. |
| [22] | Hassani H, Nikan S, Shami A. Improved exploration-exploitation trade-off through adaptive prioritized experience replay [J]. Neurocomputing, 2025, 614: No.128836. |
| [23] | Wang X, Yang Z, Chen G, et al. A reinforcement learning method of solving Markov decision processes: an adaptive exploration model based on temporal difference error [J]. Electronics, 2023, 12(19): No.4176. |
| [24] | Flennerhag S, Wang J X, Sprechmann P, et al. Temporal difference uncertainties as a signal for exploration[PP/OL]. V2. arXiv (2021-07-01) [2025-08-28].. |
| [1] | Juan CHEN, Yujie CHEN, Zongling WU, Di TIAN, Jie ZHONG. User-centric satellite edge computing architecture for task offloading optimization [J]. Journal of Computer Applications, 2026, 46(8): 2524-2532. |
| [2] | Yuanhang HUANG, Na RONG. Model-free photovoltaic hosting capacity assessment method using attention mechanism and deep reinforcement learning [J]. Journal of Computer Applications, 2026, 46(8): 2708-2715. |
| [3] | Pengyu CHEN, Baojun TIAN, Lichang ZHAO, Jiandong FANG. Learning path recommendation model integrating multi-behavior modeling and reinforcement learning [J]. Journal of Computer Applications, 2026, 46(8): 2485-2493. |
| [4] | Yanyang LIANG, Wenxuan XIE, Wei CUI, Hongfei LYU, Da LI, Dongzhou ZHONG. Robotic end-to-end dynamic grasping method based on curriculum reinforcement learning [J]. Journal of Computer Applications, 2026, 46(7): 2307-2317. |
| [5] | Xiayu WU, Hong ZHANG. Review of evolution and changes in crowd evacuation calculation methods [J]. Journal of Computer Applications, 2026, 46(4): 1309-1322. |
| [6] | Tianyu XUE, Aiping LI, Liguo DUAN. Vehicular edge computing scheme with task offloading and resource optimization [J]. Journal of Computer Applications, 2025, 45(6): 1766-1775. |
| [7] | Pengcheng XU, Lei HE, Chuan LI, Weiqi QIAN, Tun ZHAO. Deep symbolic regression method based on Transformer [J]. Journal of Computer Applications, 2025, 45(5): 1455-1463. |
| [8] | Jing WANG, Xuming FANG. Intelligent joint power and channel allocation algorithm for Wi-Fi7 multi-link integrated communication and sensing [J]. Journal of Computer Applications, 2025, 45(2): 563-570. |
| [9] | Huahua WANG, Liang HUANG, Jiajie CHEN, Jiening FANG. Dynamic allocation algorithm for multi-beam subcarriers of low orbit satellites based on deep reinforcement learning [J]. Journal of Computer Applications, 2025, 45(2): 571-577. |
| [10] | Jun ZENG, Yinghua TONG, Defang WANG. Anomaly detection method based on cumulative probability fluctuation and automated clustering [J]. Journal of Computer Applications, 2025, 45(12): 3864-3871. |
| [11] | Haoxiang XU, Dunhui YU, Yichen DENG, Kui XIAO. Knowledge graph constrained question answering model based on hierarchical reinforcement learning [J]. Journal of Computer Applications, 2025, 45(12): 3764-3770. |
| [12] | Chengyi WANG, Lei XU, Jinyin CHEN, Hongjun QIU. Cyber anti-mapping method based on adaptive perturbation [J]. Journal of Computer Applications, 2025, 45(12): 3896-3908. |
| [13] | Xiaojuan CHEN, Wei ZHANG. Task allocation of unmanned aerial vehicle for rural last-mile delivery based on reinforcement learning [J]. Journal of Computer Applications, 2025, 45(12): 4055-4063. |
| [14] | Lin WEI, Shihao ZHANG, Mengyang HE. Workflow task optimization and energy-efficient offloading method for computing power network [J]. Journal of Computer Applications, 2025, 45(12): 3916-3924. |
| [15] | Xiang KUANG, Zhen MA, Wanchun ZHU, Zhi ZHANG, Yunfei CUI. Secure and reliable service function chain deployment based on encoder-decoder structured reinforcement learning [J]. Journal of Computer Applications, 2025, 45(12): 3947-3956. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||