| [1] |
Wang J, Karatzoglou A, Arapakis I, et al. Reinforcement learning-based recommender systems with large language models for state reward and action modeling [C]// SIGIR 2024. New York: ACM, 2024: 375-385.
|
| [2] |
代珊珊,刘全. 基于动作约束深度强化学习的安全自动驾驶方法[J]. 计算机科学, 2021, 48(9): 235-243.
|
|
Dai Shanshan, Liu Quan. Action constrained deep reinforcement learning based safe automatic driving method [J]. Computer Science, 2021, 48(9): 235-243.
|
| [3] |
Wang X, Zhang J, Hou D, et al. Autonomous driving based on approximate safe action [J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(12): 14320-14328.
|
| [4] |
何浩东,符浩,王强,等. 基于深度强化学习的多机器人路径跟随与编队[J]. 计算机应用, 2024, 44(8): 2626-2633.
|
|
He Haodong, Fu Hao, Wang Qiang, et al. Multi-robot path following and formation based on deep reinforcement learning[J]. Journal of Computer Applications, 2024, 44(8): 2626-2633.
|
| [5] |
马天,席润韬,吕佳豪,等. 基于深度强化学习的移动机器人三维路径规划方法[J]. 计算机应用, 2024, 44(7): 2055-2064.
|
|
Ma Tian, Xi Runtao, Jiahao Lyu, et al. Mobile robot 3D space path planning method based on deep reinforcement learning [J]. Journal of Computer Applications, 2024, 44(7): 2055-2064.
|
| [6] |
Haarnoja T, Tang H, Abbeel P, et al. Reinforcement learning with deep energy-based policies [J]. Proceedings of Machine Learning Research, 2017, 70: 1352-1361.
|
| [7] |
Ziebart B D, Maas A, Bagnell J A, et al. Maximum entropy inverse reinforcement learning [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2008, 3: 1433-1438.
|
| [8] |
Boularias A, Kober J, Peters J. Relative entropy inverse reinforcement learning [J]. Proceedings of Machine Learning Research, 2011, 15: 182-189.
|
| [9] |
Ng A Y, Russell S J. Algorithms for inverse reinforcement learning[C]// ICML 2000. San Francisco: Morgan Kaufmann Publishers Inc., 2000: 663-670.
|
| [10] |
Nachum O, Norouzi M, Xu K, et al. Bridging the gap between value and policy based reinforcement learning [C]// NeurIPS 2017. Red Hook: Curran Associates Inc., 2017: 2772-2782.
|
| [11] |
Mnih V, Badia A P, Mirza M, et al. Asynchronous methods for deep reinforcement learning [J]. Proceedings of Machine Learning Research, 2016, 48: 1928-1937.
|
| [12] |
Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor [J]. Proceedings of Machine Learning Research, 2018, 80: 1861-1870.
|
| [13] |
Lillicrap T P, Hunt J J, Pritzel A, et al. Continuous control with deep reinforcement learning[PP/OL]. V6. arXiv (2019-07-05) [2025-08-13]..
|
| [14] |
Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms[PP/OL]. V2. arXiv (2017-08-28) [2025-08-26]..
|
| [15] |
Fujimoto S, van Hoof H, Meger D. Addressing function approximation error in actor-critic methods [J]. Proceedings of Machine Learning Research, 2018, 80: 1587-1596.
|
| [16] |
Todorov E, Erez T, Tassa Y. MuJoCo: a physics engine for model-based control [C]// IROS 2012. Piscataway: IEEE, 2012: 5026-5033.
|
| [17] |
Haarnoja T, Zhou A, Hartikainen K, et al. Soft actor-critic algorithms and applications[PP/OL]. V2. arXiv (2019-01-29) [2025-08-17]..
|
| [18] |
Igoe C, Pande S, Venkatraman S, et al. Multi-alpha soft actor-critic: overcoming stochastic biases in maximum entropy reinforcement learning [C]// ICRA 2023. Piscataway: IEEE, 2023: 7162-7168.
|
| [19] |
Wang Y, Ni T. Meta-SAC: auto-tune the entropy temperature of soft actor-critic via metagradient[PP/OL]. V2. arXiv (2020-07-31) [2025-07-17]..
|
| [20] |
Sutton R S. Learning to predict by the methods of temporal differences [J]. Machine Learning, 1988, 3(1): 9-44.
|
| [21] |
Ding Z, Huang Y, Yuan H, et al. Introduction to reinforcement learning [M]// Dong H, Ding Z, Zhang S. Deep reinforcement learning: fundamentals, research and applications. Singapore: Springer, 2020: 47-123.
|
| [22] |
Hassani H, Nikan S, Shami A. Improved exploration-exploitation trade-off through adaptive prioritized experience replay [J]. Neurocomputing, 2025, 614: No.128836.
|
| [23] |
Wang X, Yang Z, Chen G, et al. A reinforcement learning method of solving Markov decision processes: an adaptive exploration model based on temporal difference error [J]. Electronics, 2023, 12(19): No.4176.
|
| [24] |
Flennerhag S, Wang J X, Sprechmann P, et al. Temporal difference uncertainties as a signal for exploration[PP/OL]. V2. arXiv (2021-07-01) [2025-08-28]..
|