Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2800-2808.DOI: 10.11772/j.issn.1001-9081.2025081035

• Artificial intelligence • Previous Articles    

Incremental semantic mapping system based on visual SLAM

Wei CUI1, Ping ZHANG1(), Zhuoling XIAO2, Zhuohang CHEN3   

  1. 1.Digital Development Center,Luzhou Laojiao Company Limited,Luzhou Sichuan 646000,China
    2.School of Information and Communication Engineering,University of Electronic Science and Technology of China,Chengdu Sichuan 611731,China
    3.Glasgow College,University of Electronic Science and Technology of China,Chengdu Sichuan 611731,China
  • Received:2025-09-09 Revised:2026-03-18 Accepted:2026-03-20 Online:2026-04-22 Published:2026-09-10
  • Contact: Ping ZHANG
  • About author:CUI Wei, born in 1983, M. S. candidate, senior engineer. His research interests include artificial intelligence, internet of things.
    ZHANG Ping, born in 1986, senior engineer. His research interests include parallel computing, big data applications.
    XIAO Zhuoling, born in 1985, Ph. D., professor. His research interests include intelligent signal processing, internet of things.
    CHEN Zhuohang, born in 2005. His research interests include embedded development, artificial intelligence.
  • Supported by:
    General Program of National Natural Science Foundation of China(61973056)

基于视觉SLAM的增量式语义建图系统

崔伟1, 张平1(), 肖卓凌2, 陈卓航3   

  1. 1.泸州老窖股份有限公司 数字化发展中心,四川 泸州 646000
    2.电子科技大学 信息与通信工程学院,成都 611731
    3.电子科技大学 格拉斯哥学院,成都 611731
  • 通讯作者: 张平
  • 作者简介:崔伟(1983—),男,四川自贡人,高级工程师,硕士研究生,主要研究方向:人工智能、物联网
    张平(1986—),男,四川泸州人,高级工程师,主要研究方向:并行计算、大数据应用
    肖卓凌(1985—),男,四川成都人,教授,博士,主要研究方向:智能信号处理、物联网
    陈卓航(2005—),男,四川成都人,主要研究方向:嵌入式开发、人工智能。
  • 基金资助:
    国家自然科学基金面上项目(61973056)

Abstract:

Traditional semantic Simultaneous Localization and Mapping (SLAM) systems rely on training with full data, making it difficult to recognize novel categories in open environments. In addition, the systems only construct sparse maps, resulting in limited semantic representation capabilities and insufficient support for complex human-machine interaction. Therefore, an incremental semantic mapping system based on visual SLAM was proposed. In the system, a dense mapping thread was introduced into the ORB-SLAM3 (Oriented FAST and Rotated BRIEF SLAM 3.0) framework, and a newly designed semantic segmentation model, DAISS (Depth Attention Incremental Semantic Segmentation), was integrated, so as to achieve simultaneous localization and semantic mapping. Among which, a dual-branch encoder and an attention distillation mechanism were adopted by the DAISS model to improve segmentation accuracy for old and new categories. Experimental results on NYU Depth V2 dataset show that this system achieves 61.0% and 47.5% in Mean Accuracy (MA) and Mean Intersection over Union (MIoU) metrics, representing improvements of 60.9% and 48.9%, respectively, over the SATS (Self-Attention Transfer for continual Semantic segmentation) model. These results validate its strong incremental learning capability and semantic perception performance.

Key words: semantic mapping, 3D reconstruction, incremental learning, multimodal data, data fusion, deep learning

摘要:

传统语义同时定位与建图(SLAM)系统依赖全量数据进行训练,难以适应开放环境中的新类别识别,且仅构建稀疏地图,语义表达能力有限,难以支持复杂的人机交互。因此,提出一种基于视觉SLAM的增量式语义建图系统。该系统在ORB-SLAM3(Oriented FAST and Rotated BRIEF SLAM 3.0)框架中引入稠密建图线程,并融合所设计的DAISS(Depth Attention Incremental Semantic Segmentation)语义分割模型,实现同步定位与语义建图。其中,DAISS模型采用双分支编码器和注意力蒸馏机制,以提升新旧类别分割精度。在NYU Depth V2数据集上的实验结果显示,该系统的平均精度(MA)、平均交并比(MIoU)分别达到了61.0%和47.5%,相较于SATS(Self-Attention Transfer for continual Semantic segmentation)模型,分别提高60.9%和48.9%,具备良好的增量学习能力和语义感知性能。

关键词: 语义建图, 三维重建, 增量式学习, 多模态数据, 数据融合, 深度学习

CLC Number: