《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (8): 2603-2611.DOI: 10.11772/j.issn.1001-9081.2025070890

• 多媒体计算与计算机仿真 • 上一篇    下一篇

多尺度特征优化的单目深度估计框架FCMdepth

刘凤春1,2,3,4,5,6, 邵馨莹7, 张春英1,3,4,5,6(), 王立亚1,3,4,5,6, 任静1,3,4,5,6   

  1. 1.华北理工大学 理学院,河北 唐山 063210
    2.华北理工大学 轻工学院,河北 唐山 063210
    3.铁矿石优选与铁前工艺智能化河北省工程研究中心(华北理工大学),河北 唐山 063210
    4.河北省数据科学与应用重点实验室(华北理工大学),河北 唐山 063210
    5.华北理工大学 唐山市工程计算重点实验室,河北 唐山 063210
    6.华北理工大学 唐山市智能工业与图像处理技术创新中心,河北 唐山 063210
    7.华北理工大学 人工智能学院,河北 唐山 063210
  • 收稿日期:2025-08-05 修回日期:2025-10-15 接受日期:2025-10-15 发布日期:2025-11-05 出版日期:2026-08-10
  • 通讯作者: 张春英
  • 作者简介:刘凤春(1976—),男,辽宁丹东人,教授,硕士,CCF会员,主要研究方向:深度学习、数据挖掘、大数据安全、隐私保护
    邵馨莹(1999—),女,河北迁安人,硕士研究生,主要研究方向:图像处理
    张春英(1969—),女,河北迁安人,教授,博士,CCF会员,主要研究方向:深度学习、计算机视觉
    王立亚(1987—),女,河北唐山人,副教授,硕士,主要研究方向:机器学习、大数据分析
    任静(1995—),女,河南焦作人,讲师,硕士,主要研究方向:数据挖掘、机器学习、粗糙集。
  • 基金资助:
    省属高校基本科研业务费资助项目(JJC2024036);唐山市科技计划项目(24140202C)

FCMdepth: monocular depth estimation framework with multi-scale feature optimization

Fengchun LIU1,2,3,4,5,6, Xinying SHAO7, Chunying ZHANG1,3,4,5,6(), Liya WANG1,3,4,5,6, Jing REN1,3,4,5,6   

  1. 1.College of Sciences,North China University of Science and Technology,Tangshan Hebei 063210,China
    2.Qing Gong College,North China University of Science and Technology,Tangshan Hebei 063210,China
    3.Hebei Engineering Research Center for the Intelligentization of Iron Ore Optimization and Ironmaking Raw Materials Preparation Processes (North China University of Science and Technology),Tangshan Hebei 063210,China
    4.Hebei Provincial Key Laboratory of Data Science and Application (North China University of Science and Technology),Tangshan Hebei 063210,China
    5.Tangshan Key Laboratory of Engineering Computing,North China University of Science and Technology,Tangshan Hebei 063210,China
    6.Tangshan Intelligent Industry and Image Processing Technology Innovation Center,North China University of Science and Technology,Tangshan Hebei 063210,China
    7.College of Artificial Intelligence,North China University of Science and Technology,Tangshan Hebei 063210,China
  • Received:2025-08-05 Revised:2025-10-15 Accepted:2025-10-15 Online:2025-11-05 Published:2026-08-10
  • Contact: Chunying ZHANG
  • About author:LIU Fengchun, born in 1976, M. S., professor. His research interests include deep learning, data mining, big data security, privacy protection.
    SHAO Xinying, born in 1999, M. S. candidate. Her research interests include image processing.
    WANG Liya, born in 1987, M. S., associate professor. Her research interests include machine learning, big data analysis.
    REN Jing, born in 1995, M. S., lecturer. Her research interests include data mining, machine learning, rough set.
  • Supported by:
    Fundamental Research Funds for Provincial Universities(JJC2024036);Tangshan Science and Technology Program(24140202C)

摘要:

针对单目深度估计中特征提取不足和上下文建模不充分的问题,提出一种融合多尺度特征的单目深度估计框架FCMdepth,以提升预测性能。FCMdepth采用编码器-解码器(ED)结构,编码器FC-Net由MobileNetV3-F与CDBlock组成,通过多尺度特征及空洞卷积优化特征;解码器LapMA-Net结合拉普拉斯金字塔与高效多尺度注意力(EMA)模块,以增强跨尺度特征融合,并输出准确深度图。在KITTI数据集上的实验结果表明,FCMdepth相较于Lapdepth,均方根误差(RMSE)、对数均方根误差(Log_RMSE)和平方相对误差(Sq_Rel)这3项误差指标分别低0.831、0.009和0.145,3项准确率指标的均值分别提高0.4、0.8和0.3个百分点。可见,FCMdepth在多数指标上优于对比模型,为单目深度估计和复杂场景的三维重建提供了有效参考。

关键词: 单目深度估计, 多尺度注意力机制, 深度学习, 编码器-解码器结构, 多尺度特征

Abstract:

To address the issues of insufficient feature extraction and inadequate context modeling in monocular depth estimation, we proposed a multi-scale feature fused monocular depth estimation optimization framework, FCMdepth, to enhance prediction performance. FCMdepth adopted an Encoder-Decoder (ED) structure, where the encoder, FC-Net, consists of MobileNetV3-F and CDBlock, and optimized features through multi-scale features and dilated convolutions, while the decoder, LapMA-Net, combined the Laplacian pyramid with an Efficient Multi-scale Attention (EMA) module to enhance cross-scale feature fusion and output accurate depth maps. Experimental results on the KITTI datasets show that, compared to Lapdepth, FCMdepth achieves lower values for the three error metrics: Root Mean Square Error (RMSE), Root Mean Square Logarithmic Error (Log_RMSE),and Square Relative error (Sq_Rel), with reductions of 0.831, 0.009, and 0.145, respectively, and improves three accuracy metrics by of 0.4, 0.8, and 0.3 percentage points, respectively. It can be seen that FCMdepth has superior performance compared to other models in most metrics and provides an effective reference for monocular depth estimation and 3D reconstruction in complex scenes.

Key words: monocular depth estimation, multi-scale attention mechanism, deep learning, encoder-decoder structure, multi-scale feature

中图分类号: