《计算机应用》唯一官方网站

• •    下一篇

基于跨域泛化增强的立体匹配方法

惠康华1,于慕涵1,张智1,赵敏2   

  1. 1.中国民航大学 计算机科学与技术学院 2.中国民航信息网络股份有限公司
  • 收稿日期:2026-04-09 修回日期:2026-07-02 发布日期:2026-08-06 出版日期:2026-08-06
  • 通讯作者: 于慕涵
  • 作者简介: 惠康华(1982—),男,江苏连云港人,副教授,博士,主要研究方向:图像处理;于慕涵(2002—),男,天津人,硕士研究生,主要研究方向:图像处理;张智(1993—),男,山西大同人,实验师,硕士,主要研究方向:计算机视觉、图像处理;赵敏(1979—),女,山东淄博人,高级工程师,硕士,主要研究方向:航空数字化。
  • 基金资助:
    天津市自然基金联合基金重点项目(25JCLZJC00550)

Stereo matching method based on cross-domain generalization enhancement#br#
#br#

HUI Kanghua1, YU Muhan1, ZHANG Zhi1, ZHAO Min2   

  1. 1.College of Computer Science and Technology, Civil Aviation University of China 2.TravelSky Technology Limited
  • Received:2026-04-09 Revised:2026-07-02 Online:2026-08-06 Published:2026-08-06
  • About author:HUI Kanghua, born in 1982, Ph. D., associate professor. His research interests include image processing. YU Muhan, born in 2002, M. S. candidate. His research interests include image processing. ZHANG Zhi, born in 1993, M. S., Experimentalist. His research interests include computer vision, image processing. ZHAO Min, born in 1979, M. S., Senior Engineer. Her research interests include aviation digitization.
  • Supported by:
    Joint Funds of the Natural Science Foundation of Tianjin (25JCLZJC00550)

摘要: 针对立体匹配方法的跨域精度下降问题,提出一种深度立体匹配模型SFENet(Spatial noise Feature Excitation stereo matching Network)。首先,利用分层视觉转换(HVT)方法离散化训练域特征分布并在像素级视觉变换层引入空间相关性噪声以模拟真实环境的图像特征分布,提升模型面对不同目标域时的适应能力。其次,设计了一种双维特征激励机制,在构建代价体时从空间和通道两个维度增强特征,并对组相关代价体实施特征引导的激励聚合,以主动筛选和增强有效特征,抑制无关信息并保留传播更多有效信息,两者共同引导模型学习泛化性更强的域不变特征。最后,通过边缘制导视差上采样模块增强细节保持能力和全局一致性,进一步提高模型预测精度。实验结果表明,相较于基准模型IGEV(Iterative Geometry Encoding Volume),SFENet在多个跨域测试数据集上均取得更好表现,在KITTI2012数据集上,SFENet以0.1 s的推理速度代价换取3 px错误率下降6.25%,在主流立体匹配方法中达到了精度与速度上的竞争力。

关键词: 立体匹配, 域不变特征, 跨域泛化, 特征分布模拟, 代价体构建, 视差上采样

Abstract: To address the issue of cross-domain accuracy degradation in stereo matching methods, a deep stereo matching model, SFENet(Spatial noise Feature Excitation stereo matching Network), is proposed. First, a Hierarchical Visual Transformation (HVT) method is used to discretize the training domain feature distribution and spatial correlation noise is introduced at the  pixel visual transformation layer to simulate the image feature distribution of the real environment, improving the model's adaptability to different target domains. Second, a dual-dimension feature excitation mechanism is designed, enhancing features from both spatial and channel dimensions during cost volume construction, and performing feature-guided excitation aggregation on group-wise correlation volumes to actively filter and enhance effective features, suppress irrelevant information, and retain more effective information for propagation. Both mechanisms jointly guide the model to learn domain-invariant features with stronger generalization capabilities. Finally, an edge-guided disparity upsampling module enhances detail preservation and global consistency, further improving the model's prediction accuracy. Experimental results show that compared to the benchmark model IGEV(Iterative Geometry Encoding Volume), SFENet achieves better performance on multiple cross-domain test datasets. On the KITTI2012 dataset, SFENet achieves a 6.25% reduction in error rate by 3 pixels at the cost of 0.1s inference speed, achieving competitiveness in both accuracy and speed among mainstream stereo matching methods.

Key words: stereo matching, domain-invariant feature, cross-domain generalization, feature distribution simulation, cost volume construction, disparity upsampling

中图分类号: