Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2898-2909.DOI: 10.11772/j.issn.1001-9081.2025080956

• Advanced computing • Previous Articles    

Heterogeneous multi-agent reinforcement learning enabled co-optimization of UAV 3D obstacle avoidance and edge computing

Guanliang CHEN1, Yi LIU1(), Yi YU2   

  1. 1.School of Automation,Guangdong University of Technology,Guangzhou Guangdong 510006,China
    2.Changsha Electronic Industry School,Changsha Hunan 410116,China
  • Received:2025-08-19 Revised:2025-11-30 Accepted:2025-12-05 Online:2026-02-12 Published:2026-09-10
  • Contact: Yi LIU
  • About author:CHEN Guanliang, born in 1999, M. S. candidate. His research interests include edge computing.
    LIU Yi, born in 1981, Ph. D., professor. His research interests include edge computing, smart grid.
    YU Yi, born in 1984, M. S., lecturer. Her research interests include computer networks, computer graphics and image processing.
  • Supported by:
    National Key R&D Program of China(2020YFB1807805)

异构多智能体强化学习驱动的无人机三维避障与边缘计算协同优化

陈冠良1, 刘义1(), 余意2   

  1. 1.广东工业大学 自动化学院,广州 510006
    2.长沙市电子工业学校,长沙 410116
  • 通讯作者: 刘义
  • 作者简介:陈冠良(1999—),男,广东广州人,硕士研究生,主要研究方向:边缘计算
    刘义(1981—),男,广东广州人,教授,博士,主要研究方向:边缘计算、智能电网
    余意(1984—),女,湖南长沙人,讲师,硕士,主要研究方向:计算机网络、计算机图形图像处理。
  • 基金资助:
    国家重点研发计划项目(2020YFB1807805)

Abstract:

Severe challenges in terms of real-time processing and low-energy transmission brought by the development of Internet of Things (IoT) and the proliferation of mobile terminal devices are face by compute-intensive tasks. Particularly in multi-Unmanned Aerial Vehicle (UAV) -assisted Mobile Edge Computing (MEC) scenarios, communication links are constrained by obstacle blockages and UAV trajectories in complex 3D environments, and latency and energy consumption pressures are further exacerbated. For the scenario of multiple UAVs providing computational offloading services to ground users, an optimization model was established to minimize the weighted sum of the system's maximum task completion latency and total energy consumption, so as to optimize the users' discrete offloading decisions and the UAVs' continuous 3D trajectories jointly. To solve the problem of mixed (discrete-continuous) action space and strong decision coupling, a heterogeneous multi-agent algorithm UOUM (User Offloading and UAV Mobility co-optimization) was proposed. In the algorithm, under a heterogeneous multi-agent deep reinforcement learning framework, dedicated network architectures were designed for the two types of heterogeneous agents: users and UAVs. And differential reward mechanism was introduced to quantify marginal contributions of the agents, thereby solving the multi-agent credit allocation problem. Concurrently, an Artificial Potential Field (APF) was innovatively integrated as a differentiable physical constraint into the agent learning framework, so as to ensure safe obstacle avoidance for the UAVs. Simulation results show that compared to three benchmark methods (Only User Offloading optimization (OUO), Only UAV Trajectory optimization (OUT), and a standard heterogeneous multi-agent reinforcement learning (H-MARL) using only a global reward mechanism)), UOUM has advantages in various scenarios with different numbers of users, UAVs, and obstacle densities. Compared with H-MARL, UOUM has the final convergence reward improved by approximately 28.6% on average, and achieves strong environmental adaptability in terms of latency control, energy optimization, and safe obstacle avoidance.

Key words: Mobile Edge Computing (MEC), multi-agent reinforcement learning, Unmanned Aerial Vehicle (UAV) obstacle avoidance, task offloading, 3D trajectory optimization

摘要:

随着物联网(IoT)的发展与移动终端设备的激增,计算密集型任务在实时处理与低能耗传输方面面临严峻挑战。尤其在多无人机(UAV)辅助的移动边缘计算(MEC)场景中,复杂三维环境中的通信链路受障碍物遮挡与UAV轨迹限制,进一步增大了时延与能耗压力。本文针对多UAV为地面用户提供计算卸载服务的场景,建立最小化系统最大任务完成时延与总能耗加权和的优化模型,以联合优化用户的离散卸载决策与UAV的连续三维轨迹。为了解决混合(离散-连续)动作空间与强决策耦合的问题,提出异构多智能体算法UOUM(User Offloading and UAV Mobility co-optimization)。该算法在异构多智能体深度强化学习框架下,为用户和UAV两类异构智能体设计专属网络架构;引入差分奖励机制量化智能体的边际贡献,以解决多智能体的信用分配问题;同时,将人工势能场(APF)作为可微物理约束融入智能体学习框架,以确保UAV在复杂环境中的安全避障。仿真实验结果表明,与3种基准方法(仅用户卸载优化(OUO)、仅UAV轨迹优化(OUT)以及仅采用全局奖励机制的标准异构多智能体强化学习(H-MARL))相比,UOUM在不同用户数、UAV数及障碍物密度场景下均展现出优势,UOUM的最终收敛奖励比H-MARL平均提升了28.6%,且在时延控制、能耗优化和安全避障方面均展现出强大的环境适应性。

关键词: 移动边缘计算, 多智能体强化学习, 无人机避障, 任务卸载, 三维轨迹优化

CLC Number: