Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2800-2808.DOI: 10.11772/j.issn.1001-9081.2025081035
• Artificial intelligence • Previous Articles
Wei CUI1, Ping ZHANG1(
), Zhuoling XIAO2, Zhuohang CHEN3
Received:2025-09-09
Revised:2026-03-18
Accepted:2026-03-20
Online:2026-04-22
Published:2026-09-10
Contact:
Ping ZHANG
About author:CUI Wei, born in 1983, M. S. candidate, senior engineer. His research interests include artificial intelligence, internet of things.Supported by:通讯作者:
张平
作者简介:崔伟(1983—),男,四川自贡人,高级工程师,硕士研究生,主要研究方向:人工智能、物联网基金资助:CLC Number:
Wei CUI, Ping ZHANG, Zhuoling XIAO, Zhuohang CHEN. Incremental semantic mapping system based on visual SLAM[J]. Journal of Computer Applications, 2026, 46(9): 2800-2808.
崔伟, 张平, 肖卓凌, 陈卓航. 基于视觉SLAM的增量式语义建图系统[J]. 《计算机应用》唯一官方网站, 2026, 46(9): 2800-2808.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025081035
| 阶段 | OA | MA | MIoU |
|---|---|---|---|
| Session_0 | 95.482 0 | 89.731 5 | 80.930 6 |
| Session_1 | 94.906 2 | 89.278 8 | 78.512 7 |
| Session_2 | 94.418 2 | 84.769 8 | 75.077 9 |
| Session_3 | 90.987 8 | 83.761 5 | 71.717 8 |
| Session_4 | 88.937 1 | 82.185 9 | 68.582 2 |
| Session_5 | 87.062 6 | 82.180 4 | 66.405 2 |
Tab. 1 Segmentation test results on PASCAL VOC 2012 dataset
| 阶段 | OA | MA | MIoU |
|---|---|---|---|
| Session_0 | 95.482 0 | 89.731 5 | 80.930 6 |
| Session_1 | 94.906 2 | 89.278 8 | 78.512 7 |
| Session_2 | 94.418 2 | 84.769 8 | 75.077 9 |
| Session_3 | 90.987 8 | 83.761 5 | 71.717 8 |
| Session_4 | 88.937 1 | 82.185 9 | 68.582 2 |
| Session_5 | 87.062 6 | 82.180 4 | 66.405 2 |
| 模型 | 模态 | MIoU | OA | MA |
|---|---|---|---|---|
| SATS[ | RGB | 31.9 | 37.9 | |
| SATS* | RGB+Depth | 24.6 | 66.0 | 28.8 |
| Dformer-S[ | RGB+Depth | 14.1 | 80.1 | 16.9 |
| HN-network[ | RGB+Depth | 33.5 | 62.9 | 56.0 |
| CompL[ | RGB+Depth | 31.6 | — | — |
| DAISS | RGB+Depth | 80.3 |
Tab. 2 Quantitative comparison of models on NYU Depth V2 dataset
| 模型 | 模态 | MIoU | OA | MA |
|---|---|---|---|---|
| SATS[ | RGB | 31.9 | 37.9 | |
| SATS* | RGB+Depth | 24.6 | 66.0 | 28.8 |
| Dformer-S[ | RGB+Depth | 14.1 | 80.1 | 16.9 |
| HN-network[ | RGB+Depth | 33.5 | 62.9 | 56.0 |
| CompL[ | RGB+Depth | 31.6 | — | — |
| DAISS | RGB+Depth | 80.3 |
| 类别 | IoU/% | Acc/% | 类别 | IoU/% | Acc/% |
|---|---|---|---|---|---|
| wall | 56.7 | 74.3 | mirror | 43.8 | 57.0 |
| floor | 69.6 | 81.8 | sofa | 59.9 | 74.1 |
| cabinet | 77.9 | 88.0 | books | 53.4 | 73.9 |
| chair | 51.7 | 67.0 | TV | 37.5 | 51.6 |
| window | 40.2 | 52.5 | paper | 22.0 | 31.0 |
| bookshelf | 40.6 | 53.5 | towel | 25.8 | 32.9 |
| picture | 46.9 | 69.3 | ceiling | 22.9 | 29.9 |
| curtain | 8.3 | 10.5 | 平均 | 44.0 | 56.6 |
Tab. 3 Quantitative results of Session_0 semantic segmentation of 15 basic categories of SUN RGB-D dataset
| 类别 | IoU/% | Acc/% | 类别 | IoU/% | Acc/% |
|---|---|---|---|---|---|
| wall | 56.7 | 74.3 | mirror | 43.8 | 57.0 |
| floor | 69.6 | 81.8 | sofa | 59.9 | 74.1 |
| cabinet | 77.9 | 88.0 | books | 53.4 | 73.9 |
| chair | 51.7 | 67.0 | TV | 37.5 | 51.6 |
| window | 40.2 | 52.5 | paper | 22.0 | 31.0 |
| bookshelf | 40.6 | 53.5 | towel | 25.8 | 32.9 |
| picture | 46.9 | 69.3 | ceiling | 22.9 | 29.9 |
| curtain | 8.3 | 10.5 | 平均 | 44.0 | 56.6 |
| 模型 | 基础模型 | 模态 | MIoU | OA |
|---|---|---|---|---|
| FuseNet[ | CNN | RGB-D | 37.3 | 76.3 |
| D-CNN[ | CNN | RGB-D | 42.0 | — |
| SATS[ | Transformer | RGB | 35.5 | 68.2 |
| SATS* | Transformer | RGB-D | 40.4 | 74.5 |
| Dformer-L[ | Transformer | RGB-D | ||
| DAISS | Transformer+CNN | RGB-D | 44.0 | 75.7 |
Tab. 4 Quantitative comparison of models on SUN RGB-D dataset
| 模型 | 基础模型 | 模态 | MIoU | OA |
|---|---|---|---|---|
| FuseNet[ | CNN | RGB-D | 37.3 | 76.3 |
| D-CNN[ | CNN | RGB-D | 42.0 | — |
| SATS[ | Transformer | RGB | 35.5 | 68.2 |
| SATS* | Transformer | RGB-D | 40.4 | 74.5 |
| Dformer-L[ | Transformer | RGB-D | ||
| DAISS | Transformer+CNN | RGB-D | 44.0 | 75.7 |
| [1] | 黄泽霞,邵春莉. 深度学习下的视觉SLAM综述[J]. 机器人, 2023, 45(6): 756-768. |
| Huang Zexia, Shao Chunli. Survey of visual SLAM based on deep learning [J]. Robot, 2023, 45(6): 756-768. | |
| [2] | Wang T, Dhiman V, Atanasov N. Inverse reinforcement learning for autonomous navigation via differentiable semantic mapping and planning [J]. Autonomous Robots, 2023, 47(6): 809-830. |
| [3] | Sünderhauf N, Pham T T, Latif Y, et al. Meaningful maps with object-oriented semantic mapping [C]// IROS 2017. Piscataway: IEEE, 2017: 5079-5085. |
| [4] | Campos C, Elvira R, Gómez Rodríguez J J, et al. ORB-SLAM3: An accurate open-source library for visual, visual-inertial and multi-map SLAM [J]. IEEE Transactions on Robotics, 2021, 37(6): 1874-1890. |
| [5] | Ma L, Stückler J, Kerl C, et al. Multi-view deep learning for consistent semantic mapping with RGB-D cameras [C]// IROS 2017. Piscataway: IEEE, 2017: 598-605. |
| [6] | Su H, Maji S, Kalogerakis E, et al. Multi-view convolutional neural networks for 3D shape recognition [C]// ICCV 2015. Piscataway: IEEE, 2015: 945-953. |
| [7] | McCormac J, Handa A, Davison A, et al. SemanticFusion: dense 3D semantic mapping with convolutional neural networks [C]// ICRA 2017. Piscataway: IEEE, 2017: 4628-4635. |
| [8] | 徐陈,周怡君,罗晨. 动态场景下基于光流和实例分割的视觉SLAM方法[J]. 光学学报, 2022, 42(14): No.1415002. |
| Xu Chen, Zhou Yijun, Luo Chen. Visual SLAM method based on optical flow and instance segmentation for dynamic scenes [J]. Acta Optica Sinica, 2022, 42(14): No.1415002. | |
| [9] | 华春生,郭伟豪. 动态环境下的语义视觉SLAM算法研究[J]. 辽宁大学学报(自然科学版), 2022, 49(4): 289-297. |
| Hua Chunsheng, Guo Weihao. Research on a semantic vision SLAM algorithm in dynamic environments [J]. Journal of Liaoning University (Natural Science Edition), 2022, 49(4): 289-297. | |
| [10] | Wang H, Wang J, Agapito L. Co-SLAM: joint coordinate and sparse parametric encodings for neural real-time SLAM [C]// CVPR 2023. Piscataway: IEEE, 2023: 13293-13302. |
| [11] | Natan O, Miura J. DeepIPC: deeply integrated perception and control for an autonomous vehicle in real environments [J]. IEEE Access, 2024, 12: 49590-49601. |
| [12] | Yuan B, Zhao D. A survey on continual semantic segmentation: Theory, challenge, method and application [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 10891-10910. |
| [13] | Qiu Y, Shen Y, Sun Z, et al. SATS: self-attention transfer for continual semantic segmentation [J]. Pattern Recognition, 2023, 138: No.109383. |
| [14] | Silberman N, Hoiem D, Kohli P, et al. Indoor segmentation and support inference from RGBD images [C]// ECCV 2012, LNCS 7576. Berlin: Springer, 2012: 746-760. |
| [15] | 电子科技大学. 一种基于视觉SLAM的增量式语义分割系统: CN202411755485.5[P]. 2025-04-08. |
| University of Electronic Science and Technology of China. Incremental semantic segmentation system based on visual SLAM: CN202411755485.5[P]. 2025-04-08. | |
| [16] | Xie E, Wang W, Yu Z, et al. SegFormer: simple and efficient design for semantic segmentation with Transformers [C]// NeurIPS 2021. Red Hook: Curran Associates Inc., 2021: 12077-12090. |
| [17] | Liu Z, Mao H, Wu C Y, et al. A ConvNet for the 2020s [C]// CVPR 2022. Piscataway: IEEE, 2022: 11966-11976. |
| [18] | Khudjaev N, Tsoy R, Sharif S M A, et al. Dformer: learning efficient image restoration with perceptual guidance [C]// CVPRW 2024. Piscataway: IEEE, 2024: 6363-6372. |
| [19] | Lahoud J, Ghanem B. RGB-based semantic segmentation using self-supervised depth pre-training[PP/OL]. arXiv (2020-02-06) [2024-10-05].. |
| [20] | Kanakis M, Huang T E, Brüggemann D, et al. Composite learning for robust and effective dense predictions [C]// WACV 2023. Piscataway: IEEE, 2023: 2298-2307. |
| [21] | Hazirbas C, Ma L, Domokos C, et al. FuseNet: incorporating depth into semantic segmentation via fusion-based CNN architecture [C]// ACCV 2016, LNCS 10111. Cham: Springer, 2017: 213-228. |
| [22] | Wang W, Neumann U. Depth-aware CNN for RGB-D segmentation[C]// ECCV 2018, LNCS 11215. Cham: Springer, 2018: 144-161. |
| [1] | Shang LIU, Zhaosen TANG, Hongyue LIU, Linfang DONG, Jin ZHOU. Multimodal sentiment analysis model for missing modalities under shared semantic conditions [J]. Journal of Computer Applications, 2026, 46(9): 2761-2768. |
| [2] | Xiaojin GUO, Xuyang SUI, Kenan ZHOU. Channel estimation of STAR-RIS-assisted hybrid-field communication system based on deep learning [J]. Journal of Computer Applications, 2026, 46(8): 2548-2554. |
| [3] | Ming LIU, Dongqi SHEN, Ziyang MENG. Point cloud registration network with dual-branch multi-level feature fusion [J]. Journal of Computer Applications, 2026, 46(8): 2584-2593. |
| [4] | Fengchun LIU, Xinying SHAO, Chunying ZHANG, Liya WANG, Jing REN. FCMdepth: monocular depth estimation framework with multi-scale feature optimization [J]. Journal of Computer Applications, 2026, 46(8): 2603-2611. |
| [5] | Wei LIU, Weigang LI, Zhiqiang TIAN. Representation learning method of hierarchical rotation-invariant geometric structure for point cloud classification and segmentation [J]. Journal of Computer Applications, 2026, 46(8): 2594-2602. |
| [6] | Xinliang LIU, Yushi XU, Dubai LI, Yanzhao REN. Survey of knowledge graph-based question answering methods [J]. Journal of Computer Applications, 2026, 46(8): 2394-2410. |
| [7] | Xiuli DU, Xing GAO, Xiaoyu ZHANG, Chengsheng PAN, Qijie ZOU. Video snapshot compressive imaging reconstruction method based on dense spatio-temporal deformable attention [J]. Journal of Computer Applications, 2026, 46(7): 2288-2296. |
| [8] | Xiangyi WU, Hailiang YE, Feilong CAO. Point cloud completion method based on smooth-sharpen graph convolution [J]. Journal of Computer Applications, 2026, 46(7): 2267-2276. |
| [9] | Miaogen LING, Rui JING, Wei FANG. Survey of research on applications of explainable deep learning in tropical cyclone forecasting [J]. Journal of Computer Applications, 2026, 46(7): 2318-2326. |
| [10] | Jing LIU, Shaoze ZHAO, Xingang LIU, Haozhe NIU, Haipeng JI. Cross-condition microstructure data generation method for titanium alloys based on improved CGAN [J]. Journal of Computer Applications, 2026, 46(7): 2373-2382. |
| [11] | Meihua WANG, Jie HUANG, Wen WEN, Ruichu CAI, Peijie HUANG, Yuhong XU, Xinlong LIN. Gradient orthogonal projection based continual embedding method for dynamic knowledge graph [J]. Journal of Computer Applications, 2026, 46(6): 1776-1784. |
| [12] | Yi DU, Mingjin XU, Jiayi KONG, Liyao WANG, Chen ZHAO. Low-rank adaptive parameter-efficient fine-tuning algorithm based on YOLOv11 [J]. Journal of Computer Applications, 2026, 46(6): 1738-1745. |
| [13] | Zhenkai XIONG, Mengjun XU, Yinyin SUN, Xin WANG. Maritime ship detection algorithm under complex weather environments based on enhanced YOLOv8 [J]. Journal of Computer Applications, 2026, 46(6): 1998-2006. |
| [14] | Lili HE, Meng CAO, Lei ZHANG, Hongjun PAN, Yi LIU, Chengxin SUN. Sign language generation model based on Kolmogorov-Arnold network and diffusion Transformer [J]. Journal of Computer Applications, 2026, 46(6): 1801-1810. |
| [15] | Xinyao LIU, Jun LIANG, Jiahao LONG, Renliang YAN. Fine-grained Chinese herbal medicine image classification based on feature fusion and channel information compensation [J]. Journal of Computer Applications, 2026, 46(5): 1677-1683. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||