《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (8): 2652-2659.DOI: 10.11772/j.issn.1001-9081.2025070856

• 多媒体计算与计算机仿真 • 上一篇    

基于改进实时检测Transformer的遥感图像目标检测算法

黄麟1, 李英顺2(), 佟维妍1, 张树园3, 王子涵4   

  1. 1.沈阳工业大学 化工过程自动化学院,沈阳 111003
    2.大连理工大学 控制科学与工程学院,大连 辽宁 116024
    3.内蒙古一机集团北方实业有限公司,包头 014030
    4.桂林电子科技大学 计算机与信息安全学院,广西 桂林 541004
  • 收稿日期:2025-07-29 修回日期:2025-09-19 接受日期:2025-09-25 发布日期:2025-11-05 出版日期:2026-08-10
  • 通讯作者: 李英顺
  • 作者简介:黄麟(2001—),女(满族),吉林东丰人,硕士研究生,主要研究方向:计算机视觉
    李英顺(1971—),女(朝鲜族),辽宁抚顺人,教授,博士,主要研究方向:人工智能、模式识别与智能系统、故障诊断
    佟维妍(1981—),女(满族),辽宁葫芦岛人,副教授,硕士,主要研究方向:图像分割、故障诊断
    张树园(1983—),男,内蒙古包头人,工程师,主要研究方向:机器人与自动化、机械制造
    王子涵(2004—),男,山东潍坊人,主要研究方向:图形图像处理、嵌入式系统、计算机软件。
  • 基金资助:
    辽宁省科学技术计划项目(2022JH1/10400007)

Remote sensing image object detection algorithm based on improved real-time detection Transformer

Lin HUANG1, Yingshun LI2(), Weiyan TONG1, Shuyuan ZHANG3, Zihan WANG4   

  1. 1.School of Chemical Process Automation,Shenyang University of Technology,Shenyang Liaoning 111003,China
    2.School of Control Science and Engineering,Dalian University of Technology,Dalian Liaoning 116024,China
    3.Inner Mongolia First Machinery Group Northern Industrial Company Limited,Baotou Inner Mongolia 014030,China
    4.School of Computer and Information Security,Guilin University of Electronic Technology,Guilin Guangxi 541004,China
  • Received:2025-07-29 Revised:2025-09-19 Accepted:2025-09-25 Online:2025-11-05 Published:2026-08-10
  • Contact: Yingshun LI
  • About author:HUANG Lin, born in 2001, M. S. candidate. Her research interests include computer vision.
    TONG Weiyan, born in 1981, M. S., associate professor. Her research interests include image segmentation, fault diagnosis.
    ZHANG Shuyuan, born in 1983, engineer. His research interests include robotics and automation, mechanical manufacturing.
    WANG Zihan, born in 2004. His research interests include graphics and image processing, embedded systems development, computer software and theory.
  • Supported by:
    Liaoning Provincial Science and Technology Program(2022JH1/10400007)

摘要:

针对遥感图像背景复杂、小目标众多且存在长尾型目标导致的检测精度低的问题,提出一种基于改进实时检测Transformer (RT-DETR)的遥感图像目标检测算法SLT-DETR (Small and Long-Tailed object-aware real-time DEtection TRansformer)。首先,采用多尺度特征增强金字塔(MFEP)优化原有的连续卷积特征融合,并引入OmniKernel进一步增强对小目标的检测能力;其次,引入内容引导的注意力对P3、P4和P5的输出特征进行融合,从而有效解决信息分布失衡问题;最后,使用FasterBlock模块实现模型的轻量化,并提升推理速度。实验结果表明,在自建遥感数据集上,SLT-DETR的平均精度均值(mAP)相较于RT-DETR提升了5.0%,模型参数量降低了1.9×106;相较于经典的YOLOv8m、YOLOv8l以及Swin Transformer模型,SLT-DETR在mAP上分别提高了5.7%、3.7%和10.9%。在RSOD和HIT-UAV公开数据集上,SLT-DETR在mAP上分别提高了1.1%和1.3%。可见,所提模型能够在保证轻量化的同时有效提升对小目标与长尾型目标的检测能力。

关键词: 目标检测, 实时检测Transformer, 多尺度特征融合, 轻量化, 倒残差结构

Abstract:

To address low detection precision of remote sensing image object detection caused by complex backgrounds, numerous small objects, and long-tail objects, a remote sensing image object detection algorithm named SLT-DETR (Small and Long-Tailed object-aware real-time DEtection TRansformer) was developed on the basis of improved Real-Time DEtection TRansformer (RT-DETR). First, a Multi-scale Feature Enhancement Pyramid (MFEP) was employed to optimize the original continuous convolutional fusion, and an OmniKernel module was introduced to further enhance the ability of small object detection. Then, the content-guided attention was utilized to fuse features output by P3, P4, and P5, thereby the addressing information distribution imbalance effectively. Finally, a FasterBlock module was adopted for lightweighting model with improved inference speed. Experimental results show that on the self-constructed remote sensing dataset, SLT-DETR improves the mean Average Precision (mAP) by 5.0% compared to RT-DETR, with the number of model parameters reduced by 1.9×106; compared with classic models YOLOv8m, YOLOv8l, and Swin Transformer, SLT-DETR achieves mAP improvements of 5.7%, 3.7%, and 10.9%, respectively; on the RSOD and HIT-UAV datasets, SLT-DETR achieves improvements in mAP of 1.1% and 1.3%, respectively. It can be seen that the proposed model enhances the detection of small and long-tailed objects effectively while maintaining lightweight model design.

Key words: object detection, Real-Time DEtection TRansformer (RT-DETR), multi-scale feature fusion, lightweighting, inverted residual structure

中图分类号: