Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (7): 2307-2317.DOI: 10.11772/j.issn.1001-9081.2025060749

• Multimedia computing and computer simulation • Previous Articles    

Robotic end-to-end dynamic grasping method based on curriculum reinforcement learning

Yanyang LIANG1, Wenxuan XIE2(), Wei CUI2, Hongfei LYU2, Da LI2, Dongzhou ZHONG1   

  1. 1.School of Electronics and Information Engineering,Wuyi University,Jiangmen Guangdong 529020,China
    2.School of Mechanical and Automation Engineering,Wuyi University,Jiangmen Guangdong 529020,China
  • Received:2025-07-09 Revised:2025-09-12 Accepted:2025-09-24 Online:2025-10-27 Published:2026-07-10
  • Contact: Wenxuan XIE
  • About author:LIANG Yanyang, born in 1980, Ph. D., associate professor. His research interests include machine vision, robotic motion control.
    CUI Wei, born in 2000, M. S. candidate. His research interests include machine vision, image processing.
    LYU Hongfei, born in 1998, M. S. candidate. Her research interests include robotic motion control, machine learning.
    LI Da, born in 2002, M. S. candidate. His research interests include embodied control.
    ZHONG Dongzhou, born in 1977, Ph. D., professor. His research interests include laser and optoelectronics, photonic neural networks.
  • Supported by:
    National Natural Science Foundation of China(62075168)

基于课程强化学习的机器人端到端动态抓取方法

梁艳阳1, 谢文轩2(), 崔伟2, 吕洪妃2, 李达2, 钟东洲1   

  1. 1.五邑大学 电子与信息工程学院,广东 江门 529020
    2.五邑大学 机械与自动化工程学院,广东 江门 529020
  • 通讯作者: 谢文轩
  • 作者简介:梁艳阳(1980—),男,广东阳江人,副教授,博士,主要研究方向:机器视觉、机器人运动控制
    崔伟(2000—),男,河南周口人,硕士研究生,主要研究方向:机器视觉、图像处理
    吕洪妃(1998—),女,河南新乡人,硕士研究生,主要研究方向:机器人运动控制、机器学习
    李达(2002—),男,湖南岳阳人,硕士研究生,主要研究方向:具身控制
    钟东洲(1977—),男,江西瑞金人,教授,博士,主要研究方向:激光与光电子、光子神经网络。
  • 基金资助:
    国家自然科学基金资助项目(62075168)

Abstract:

To address the low efficiency and the difficulty of balancing high success rate with motion smoothness in robotic end-to-end dynamic grasping tasks, a robotic end-to-end dynamic grasping method based on Curriculum Reinforcement Learning (CRL) was proposed. First, a multi-modal input network that fuses color images, depth maps, and robot proprioceptive states was constructed to map raw sensory data directly to continuous action commands for the end-effector. Second, a curriculum mechanism with synchronously increasing difficulty and smoothness constraints was designed, and combined with a staged reward function, the agent was guided to master grasping from static to dynamic ones progressively. Finally, Domain Randomization (DR) was employed to enhance the policy's transfer capability of Simulation-to-Reality (Sim-to-Real). Simulation results show that the proposed method achieves a grasping success rate of nearly 100% at target speeds ranging from 0.15 to 0.40 m/s, elevating the upper speed limit for stable grasping from 0.25 m/s of the object detection-based baseline method to 0.40 m/s. Compared to Simple Curriculum Learning (SimpleCL) with only increasing difficulty, the proposed method increases the success rate by 3.6 percentage points and reduces the average joint acceleration and jerk norm by 58.34% and 69.25%, respectively, in the most difficult test. In physical experiments, the grasping success rates of the proposed method for static scene and two dynamic scenes are 95.0%, 90.0%, and 70.0%, respectively. It can be seen that this method effectively coordinates success rate and smoothness in robotic dynamic grasping tasks by jointly optimizing task difficulty and behavioral constraints.

Key words: dynamic grasping, reinforcement learning, curriculum learning, end-to-end learning, robotic manipulation

摘要:

针对机器人学习端到端动态抓取任务时效率低且难以兼顾高成功率与运动平滑性的问题,提出一种基于课程强化学习(CRL)的机器人端到端动态抓取方法。首先,构建融合彩色图、深度图与机器人本体状态的多模态输入网络,将原始感知直接映射为末端执行器的连续动作指令;其次,设计难度与平滑性约束同步递增的课程机制,并结合阶段化的奖励函数引导智能体从静态抓取逐步掌握至动态抓取;最后,采用域随机化(DR)增强策略从仿真到现实(Sim-to-Real)的迁移能力。仿真实验结果表明,在0.15~0.40 m/s的目标速度下,本文方法的抓取成功率接近100%,将稳定抓取的速度上限从基于目标检测的基线方法的0.25 m/s提升至0.40 m/s;相较于仅有难度递增的简单课程学习(SimpleCL),本文方法在最高难度测试中的抓取成功率提升了3.6个百分点,且关节加速度与加加速度的平均范数分别降低了58.34%和69.25%;在物理实验中,本文方法对静态场景及2种动态场景的抓取成功率分别为95.0%、90.0%和70.0%。本文方法通过协同优化任务难度与行为约束,在机器人动态抓取任务中实现了成功率与平滑性的有效协调。

关键词: 动态抓取, 强化学习, 课程学习, 端到端学习, 机器人操作

CLC Number: