Journals
  Publication Years
  Keywords
Search within results Open Search
Please wait a minute...
For Selected: Toggle Thumbnails
Dynamic dilated convolution and hierarchical local-global attention model based on improved Lite-Mono architecture
Guanghui LI, Licheng QU
Journal of Computer Applications    2026, 46 (8): 2567-2576.   DOI: 10.11772/j.issn.1001-9081.2025070832
Abstract121)   HTML0)    PDF (1665KB)(6)       Save

Monocular depth estimation is a core technology for 3D environmental perception and is of significant value in applications such as autonomous driving and robot navigation. As a representative lightweight monocular depth estimation model in these fields, Lite-Mono achieves an excellent balance between performance and efficiency through dilated convolutions and a local-global attention mechanism. However, there are two major bottlenecks in current Lite-Mono architecture. The first one is that fixed dilation rate in the dilated convolutions leads to mismatch between receptive field and target scale, so that features of small objects are lost due to excessively large detailed rate, while contextual information of large objects is lacked due to insufficient receptive field. The second one is that high computational complexity of OH2W2C) in local-global attention module hinders processing of high-resolution input, thus limiting real-time application. To address these issues, an enhanced model based on Lite-Mono architecture, DDHL (Dynamic Dilated convolution and Hierarchical Local-global attention), was proposed. First, a Dynamic Dilated Convolution Module (DDCM) was introduced to adjust the receptive field in real time via a fusion weight rate predictor and generate adaptive weights by combining channel attention. Second, a hierarchical local-global attention module was designed to reduce the computational complexity to O((HW/M)2C). Finally, a learnable global token was incorporated to establish cross-window dependencies, so as to assist models in understanding and analyzing visual content more accurately in complex scenes, thereby enhancing perception and modeling capabilities of the overall scene. Experimental results on Make3D dataset demonstrate that DDHL model achieves significant improvements on key metrics: Absolute Relative Error (Abs Rel) is reduced from 0.462 to 0.290 with a decrease of 37.2%, and Squared Relative Error (Sq Rel) is decreased by 49.7%. It can be seen that this model achieves good balance between accuracy and efficiency, demonstrating practical application value.

Table and Figures | Reference | Related Articles | Metrics