Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (8): 2567-2576.DOI: 10.11772/j.issn.1001-9081.2025070832
• Multimedia computing and computer simulation • Previous Articles Next Articles
Received:2025-07-24
Revised:2025-09-25
Accepted:2025-09-25
Online:2025-11-05
Published:2026-08-10
Contact:
Licheng QU
About author:LI Guanghui, born in 2000, M. S. candidate. His research interests include big data, artificial intelligence.
Supported by:通讯作者:
屈立成
作者简介:李光辉(2000—),男,河南洛阳人,硕士研究生,主要研究方向:大数据、人工智能
基金资助:CLC Number:
Guanghui LI, Licheng QU. Dynamic dilated convolution and hierarchical local-global attention model based on improved Lite-Mono architecture[J]. Journal of Computer Applications, 2026, 46(8): 2567-2576.
李光辉, 屈立成. 基于改进Lite-Mono架构的动态空洞卷积和分层注意力模型[J]. 《计算机应用》唯一官方网站, 2026, 46(8): 2567-2576.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025070832
| 输入尺寸 | 传统交互时间/ms | Delaunay三角剖分 交互时间/ms | 效率提升倍数 |
|---|---|---|---|
| 128×128 | 18.6 | 6.4 | 2.90 |
| 256×256 | 74.5 | 26.5 | 2.81 |
| 512×512 | 298.2 | 105.1 | 2.83 |
Tab. 1 Comparison of cross-window modeling efficiency results
| 输入尺寸 | 传统交互时间/ms | Delaunay三角剖分 交互时间/ms | 效率提升倍数 |
|---|---|---|---|
| 128×128 | 18.6 | 6.4 | 2.90 |
| 256×256 | 74.5 | 26.5 | 2.81 |
| 512×512 | 298.2 | 105.1 | 2.83 |
| 模型 | 误差(↓) | 准确率(↑) | 推理时间/ms(↓) | 模型大小/MB(↓) | ||||
|---|---|---|---|---|---|---|---|---|
| Abs Rel | Sq Rel | RMSE | δ<1.25 | δ<1.252 | δ<1.253 | |||
| DepthFormer[ | 0.438 | 3.998 | 7.329 | 0.339 | 0.612 | 0.779 | 9 515.56 | 86.00 |
| GLPDepth[ | 0.446 | 4.045 | 8.420 | 0.328 | 0.572 | 0.741 | 93.37 | 23.00 |
| SPIdepth[ | 0.336 | 2.283 | 6.646 | 0.464 | 0.749 | 0.888 | 15.50 | 3 014.00 |
| MonoDiffusion[ | 0.445 | 3.360 | 7.872 | 0.37 | 0.607 | 0.791 | 16.26 | 12.00 |
| DepthFM[ | 0.441 | 11.659 | 8.967 | 0.631 | 0.751 | 0.828 | 178.11 | 3 276.00 |
| UniDepthV2[ | 0.546 | 6.771 | 8.286 | 0.353 | 0.662 | 0.829 | 31.38 | 137.00 |
| MonSter[ | 0.418 | 6.760 | 7.528 | 0.497 | 0.711 | 0.841 | 64.60 | 1 434.00 |
| Depth Pro[ | 0.440 | 3.290 | 7.808 | 0.497 | 0.711 | 0.841 | 18.28 | 3 631.00 |
| Lite-Mono[ | 0.462 | 3.950 | 8.454 | 0.346 | 0.589 | 0.767 | 86.77 | 0.90 |
| 本文模型 | 0.290 | 1.985 | 5.988 | 0.550 | 0.814 | 0.919 | 9.56 | 0.97 |
Tab. 2 Comparison experiment results
| 模型 | 误差(↓) | 准确率(↑) | 推理时间/ms(↓) | 模型大小/MB(↓) | ||||
|---|---|---|---|---|---|---|---|---|
| Abs Rel | Sq Rel | RMSE | δ<1.25 | δ<1.252 | δ<1.253 | |||
| DepthFormer[ | 0.438 | 3.998 | 7.329 | 0.339 | 0.612 | 0.779 | 9 515.56 | 86.00 |
| GLPDepth[ | 0.446 | 4.045 | 8.420 | 0.328 | 0.572 | 0.741 | 93.37 | 23.00 |
| SPIdepth[ | 0.336 | 2.283 | 6.646 | 0.464 | 0.749 | 0.888 | 15.50 | 3 014.00 |
| MonoDiffusion[ | 0.445 | 3.360 | 7.872 | 0.37 | 0.607 | 0.791 | 16.26 | 12.00 |
| DepthFM[ | 0.441 | 11.659 | 8.967 | 0.631 | 0.751 | 0.828 | 178.11 | 3 276.00 |
| UniDepthV2[ | 0.546 | 6.771 | 8.286 | 0.353 | 0.662 | 0.829 | 31.38 | 137.00 |
| MonSter[ | 0.418 | 6.760 | 7.528 | 0.497 | 0.711 | 0.841 | 64.60 | 1 434.00 |
| Depth Pro[ | 0.440 | 3.290 | 7.808 | 0.497 | 0.711 | 0.841 | 18.28 | 3 631.00 |
| Lite-Mono[ | 0.462 | 3.950 | 8.454 | 0.346 | 0.589 | 0.767 | 86.77 | 0.90 |
| 本文模型 | 0.290 | 1.985 | 5.988 | 0.550 | 0.814 | 0.919 | 9.56 | 0.97 |
| 模型 | 误差(↓) | 准确率(↑) | 运行时间/ms(↓) | 模型 大小/KB(↓) | ||||
|---|---|---|---|---|---|---|---|---|
| Abs Rel | Sq Rel | RMSE | δ<1.25 | δ<1.252 | δ<1.253 | |||
| 本文模型 | 0.290 | 1.985 | 5.988 | 0.550 | 0.814 | 0.919 | 9.56 | 994 |
| 原始模型+动态空洞卷积模块 | 0.328 | 2.246 | 6.705 | 0.459 | 0.746 | 0.886 | 14.43 | 998 |
| 原始模型+局部注意力机制 | 0.435 | 3.436 | 7.839 | 0.362 | 0.633 | 0.791 | 87.35 | 919 |
| 原始模型+全局注意力机制 | 0.402 | 2.981 | 7.358 | 0.385 | 0.675 | 0.827 | 87.62 | 916 |
| 原始模型+分层局部-全局注意力机制 | 0.371 | 2.595 | 6.966 | 0.417 | 0.712 | 0.862 | 87.98 | 912 |
| 原始模型 | 0.462 | 3.950 | 8.454 | 0.346 | 0.589 | 0.767 | 86.77 | 922 |
Tab. 3 Ablation experimental results of original model and modified modules
| 模型 | 误差(↓) | 准确率(↑) | 运行时间/ms(↓) | 模型 大小/KB(↓) | ||||
|---|---|---|---|---|---|---|---|---|
| Abs Rel | Sq Rel | RMSE | δ<1.25 | δ<1.252 | δ<1.253 | |||
| 本文模型 | 0.290 | 1.985 | 5.988 | 0.550 | 0.814 | 0.919 | 9.56 | 994 |
| 原始模型+动态空洞卷积模块 | 0.328 | 2.246 | 6.705 | 0.459 | 0.746 | 0.886 | 14.43 | 998 |
| 原始模型+局部注意力机制 | 0.435 | 3.436 | 7.839 | 0.362 | 0.633 | 0.791 | 87.35 | 919 |
| 原始模型+全局注意力机制 | 0.402 | 2.981 | 7.358 | 0.385 | 0.675 | 0.827 | 87.62 | 916 |
| 原始模型+分层局部-全局注意力机制 | 0.371 | 2.595 | 6.966 | 0.417 | 0.712 | 0.862 | 87.98 | 912 |
| 原始模型 | 0.462 | 3.950 | 8.454 | 0.346 | 0.589 | 0.767 | 86.77 | 922 |
| 分支数量 | 空洞率取值 | Abs Rel | Sq Rel | RMSE | δ<1.253 | 运行时间/ms |
|---|---|---|---|---|---|---|
| 3 | 1,3,6 | 0.290 | 1.985 | 5.988 | 0.919 | 9.56 |
| 2 | 1,6 | 0.302 | 2.114 | 6.175 | 0.893 | 12.21 |
| 4 | 1,3,6,9 | 0.297 | 2.019 | 6.053 | 0.906 | 10.57 |
Tab. 4 Ablation experiment results for different numbers of dilated branches of CDC module in DDHL model
| 分支数量 | 空洞率取值 | Abs Rel | Sq Rel | RMSE | δ<1.253 | 运行时间/ms |
|---|---|---|---|---|---|---|
| 3 | 1,3,6 | 0.290 | 1.985 | 5.988 | 0.919 | 9.56 |
| 2 | 1,6 | 0.302 | 2.114 | 6.175 | 0.893 | 12.21 |
| 4 | 1,3,6,9 | 0.297 | 2.019 | 6.053 | 0.906 | 10.57 |
| [1] | 邓慧萍,盛志超,向森,等. 基于语义导向的光场图像深度估计[J]. 电子与信息学报, 2022, 44(8): 2940-2948. |
| Deng Huiping, Sheng Zhichao, Xiang Sen, et al. Depth estimation based on semantic guidance for light field image[J]. Journal of Electronics and Information Technology, 2022, 44(8): 2940-2948. | |
| [2] | 程德强,张华强,寇旗旗,等. 基于层级特征融合的室内自监督单目深度估计[J]. 光学精密工程, 2023, 31(20): 2993-3009. |
| Cheng Deqiang, Zhang Huaqiang, Kou Qiqi, et al. Indoor self-supervised monocular depth estimation based on level feature fusion[J]. Optics and Precision Engineering, 2023, 31(20): 2993-3009. | |
| [3] | 肖磊,胡鹏,马俊杰. 局部注意力作用下基于全局信息关联的自监督单目深度估计模型[J]. 激光与光电子学进展, 2025, 62(8): No.0815010. |
| Xiao Lei, Hu Peng, Ma Junjie. Self-supervised monocular depth estimation model based on global information correlation under influence of local attention[J]. Laser and Optoelectronics Progress, 2025, 62(8): No.0815010. | |
| [4] | 熊炜,陈奕博,张丽真,等. 利用多帧序列影像的自监督单目深度估计[J]. 计算机应用, 2024, 44(12): 3907-3914. |
| Xiong Wei, Chen Yibo, Zhang Lizhen, et al. Self-supervised monocular depth estimation using multi-frame sequential images[J]. Journal of Computer Applications, 2024, 44(12): 3907-3914. | |
| [5] | Zhang N, Nex F, Vosselman G, et al. Lite-Mono: a lightweight CNN and Transformer architecture for self-supervised monocular depth estimation[C]// CVPR 2023. Piscataway: IEEE, 2023: 18537-18546. |
| [6] | Li Z, Chen Z, Liu X, et al. DepthFormer: exploiting long-range correlation and local information for accurate monocular depth estimation[J]. Machine Intelligence Research, 2023, 20(6): 837-854. |
| [7] | Bae J, Moon S, Im S. MonoFormer: towards generalization of self-supervised monocular depth estimation with Transformers[PP/OL]. V1. arXiv (2022-03-23) [2025-06-11].. |
| [8] | 江俊君,李震宇,刘贤明. 基于深度学习的单目深度估计方法综述[J]. 计算机学报, 2022, 45(6): 1276-1307. |
| Jiang Junjun, Li Zhenyu, Liu Xianming. Deep learning based monocular depth estimation:a survey[J]. Chinese Journal of Computers, 2022, 45(6): 1276-1307. | |
| [9] | 程德强,徐帅,吕晨,等. 方向感知增强的轻量级自监督单目深度估计方法[J]. 电子与信息学报, 2024, 46(9): 3683-3692. |
| Cheng Deqiang, Xu Shuai, Chen Lyu, et al. Lightweight self-supervised monocular depth estimation method with direction-aware enhancement[J]. Journal of Electronics and Information Technology, 2024, 46(9): 3683-3692. | |
| [10] | Howard A G, Zhu M, Chen B, et al. MobileNets: efficient convolutional neural networks for mobile vision applications[PP/OL]. V1. arXiv (2017-04-17) [2025-06-11].. |
| [11] | Zhang X, Zhou X, Lin M, et al. ShuffleNet: an extremely efficient convolutional neural network for mobile devices[C]// CVPR 2018. Piscataway: IEEE, 2018: 6848-6856. |
| [12] | 袁健,李佳慧. 融合先验信息的残差空间注意力人脸超分辨率重建模型[J]. 小型微型计算机系统, 2023, 44(5): 1035-1042. |
| Yuan Jian, Li Jiahui. Residual spatial attention face super resolution algorithm based on prior information-fusion[J]. Journal of Chinese Computer Systems, 2023, 44(5): 1035-1042. | |
| [13] | 高琛,冯德俊,胡金林,等. 改进特征金字塔网络的遥感影像崩滑体提取[J]. 测绘科学, 2021, 46(11): 32-38. |
| Gao Chen, Feng Dejun, Hu Jinlin, et al. Collapse and landslide extraction from remote sensing image based on improved feature pyramid network[J]. Science of Surveying and Mapping, 2021, 46(11): 32-38. | |
| [14] | 沈希忠,谢旭. 带钢表面缺陷的RepVGG网络改进及其识别[J]. 现代制造工程, 2023(5): 121-126. |
| Shen Xizhong, Xie Xu. RepVGG networks improvement of surface defects in strip steel and their identification[J]. Modern Manufacturing Engineering, 2023(5): 121-126. | |
| [15] | Saxena A, Sun M, Ng A Y. Make 3D: learning 3D scene structure from a single still image[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2009, 31(5): 824-840. |
| [16] | 梁水波,刘紫燕,孙昊堃,等. Transformer与多尺度注意力的自监督单目图像深度估计[J]. 小型微型计算机系统, 2023, 44(4): 825-831. |
| Liang Shuibo, Liu Ziyan, Sun Haokun, et al. Self-supervised monocular image depth estimation primed by Transformer and multi-scale attention scheme[J]. Journal of Chinese Computer Systems, 2023, 44(4): 825-831. | |
| [17] | Loshchilov I, Hutter F. SGDR: stochastic gradient descent with warm restarts[PP/OL]. V5. arXiv (2017-05-03) [2025-06-11].. |
| [18] | Godard C, Oisin Mac Aodha O, Firman M, et al. Digging into self-supervised monocular depth estimation[C]// ICCV 2019. Piscataway: IEEE, 2019: 3827-3837. |
| [19] | Lyu X, Liu L, Wang M, et al. HR-Depth: high resolution self-supervised monocular depth estimation[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(3): 2294-2301. |
| [20] | Zhou Z, Fan X, Shi P, et al. R-MSFM: recurrent multi-scale feature modulation for monocular depth estimating[C]// ICCV 2021. Piscataway: IEEE, 2021: 12757-12766. |
| [21] | Kim D, Ka W, Ahn P, et al. Global-local path networks for monocular depth estimation with vertical CutDepth[PP/OL]. V1. arXiv (2022-10-29) [2025-06-11].. |
| [22] | Lavreniuk M, Lavreniuk A. SPIdepth: strengthened pose information for self-supervised monocular depth estimation[C]// CVPR 2025 . Piscataway: IEEE, 2025: 865-875. |
| [23] | Shao S, Pei Z, Chen W, et al. MonoDiffusion: self-supervised monocular depth estimation using diffusion model[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(4): 3664-3678. |
| [24] | Gui M, Schusterbauer J, Prestel U, et al. DepthFM: fast generative monocular depth estimation with flow matching[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(3): 3203-3211. |
| [25] | Piccinelli L, Sakaridis C, Yang Y H, et al. UniDepthV2: universal monocular metric depth estimation made simpler[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2026, 48(3): 2354-2367. |
| [26] | Cheng J, Liu L, Xu G, et al. MonSter: marry monodepth to stereo unleashes power[C]// CVPR 2025. Piscataway: IEEE, 2025: 6273-6282. |
| [27] | Bochkovskii A, Delaunoy A, Germain H, et al. Depth Pro: sharp monocular metric depth in less than a second[PP/OL]. V2. arXiv (2025-04-21) [2025-06-11].. |
| [1] | Dong LI, Yiji ZHAO, Haiyan DING, Hao WU. Soft whitening inspired non-contrastive SSL framework for time-series forecasting [J]. Journal of Computer Applications, 2026, 46(8): 2457-2466. |
| [2] | Haoran YUAN, Huan LIU, Pengfei JIAO, Zhidong ZHAO, Xianfei ZHANG, Zunliang LIU. Masked autoencoder enhanced dynamic heterogeneous graph representation learning model [J]. Journal of Computer Applications, 2026, 46(6): 1728-1737. |
| [3] | Yu CHEN, Shuaikang QI, Liwei XU, Haotian ZHU. Social bot detection framework fusing multi-scale wavelet enhancement and self-supervised learning [J]. Journal of Computer Applications, 2026, 46(6): 1756-1766. |
| [4] | Hang QI, Tingting DONG, Yongqiang NAI, Xian MO. Contrastive collaborative filtering method based on graph diffusion generation and adaptive sampling [J]. Journal of Computer Applications, 2026, 46(6): 1818-1828. |
| [5] | Chunyong YIN, Bufan ZHANG. Multi-scale based multivariate time series anomaly detection model [J]. Journal of Computer Applications, 2026, 46(3): 790-797. |
| [6] | Wen LI, Kairong LI, Kai YANG. Subgraph-aware contrastive learning with data augmentation [J]. Journal of Computer Applications, 2026, 46(1): 1-9. |
| [7] | Chao LIU, Yanhua YU. Knowledge-aware recommendation model combining denoising strategy and multi-view contrastive learning [J]. Journal of Computer Applications, 2025, 45(9): 2827-2837. |
| [8] | Zonghang WU, Dong ZHANG, Guanyu LI. Multimodal fusion recommendation algorithm based on joint self-supervised learning [J]. Journal of Computer Applications, 2025, 45(6): 1858-1868. |
| [9] | Guangju YANG, Tianjian LUO, Kaijun WANG, Siqi YANG. Multi-branch multi-view based contextual contrastive representation learning method for time series [J]. Journal of Computer Applications, 2025, 45(4): 1042-1052. |
| [10] | Junyi ZHU, Leilei CHANG, Xiaobin XU, Zhiyong HAO, Haiyue YU, Jiang JIANG. Self-supervised learning method using minimal prior knowledge [J]. Journal of Computer Applications, 2025, 45(4): 1035-1041. |
| [11] | Jianfeng YANG, Bin CHEN, Yuxuan LI. Self-supervised point cloud anomaly detection method based on point cloud reconstruction [J]. Journal of Computer Applications, 2025, 45(10): 3302-3310. |
| [12] | Zhenyuan LIANG, Songlin JIANG, Songhao ZHU. Self-supervised image denoising based on blind-ring network and random recovery mask [J]. Journal of Computer Applications, 2025, 45(10): 3311-3319. |
| [13] | Tingjie TANG, Jiajin HUANG, Jin QIN. Session-based recommendation with graph auxiliary learning [J]. Journal of Computer Applications, 2024, 44(9): 2711-2718. |
| [14] | Tong CHEN, Fengyu YANG, Yu XIONG, Hong YAN, Fuxing QIU. Construction method of voiceprint library based on multi-scale frequency-channel attention fusion [J]. Journal of Computer Applications, 2024, 44(8): 2407-2413. |
| [15] | Jiong WANG, Taotao TANG, Caiyan JIA. PAGCL: positive augmentation graph contrastive learning recommendation method without negative sampling [J]. Journal of Computer Applications, 2024, 44(5): 1485-1492. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||
