Journal of Computer Applications

    Next Articles

Multi-scale feature-guided adaptive local queries for vehicle re-identification

  

  • Received:2026-04-20 Revised:2026-07-04 Accepted:2026-07-08 Online:2026-07-14 Published:2026-07-14

多尺度特征引导自适应局部查询的车辆重识别

王 璐,赵凯璐*,陆筱霞,熊 凯   

  1. 中原工学院 计算机学院, 郑州 450000
  • 通讯作者: 赵凯璐
  • 基金资助:
    面向车路云一体化的城市交通大场景敏捷感知方法与特征数据 高效编码关键技术研究

Abstract: To address the issues of insufficient utilization of intermediate spatial information and inadequate modeling of local discriminative features in Vehicle Re-Identification (Vehicle Re-ID) under multi-view conditions, a Multi-Scale feature-guided Adaptive Local queries (MSAL) method for vehicle re-identificationwas proposed. In this method, a Vision Transformer (ViT) was used as the backbone network to extract image patch features from different depths. A hierarchical spatial representation was constructed through a Feature Pyramid Network-based Multi-Scale Fusion (FPN-MSF) module to enhance the utilization of intermediate spatial information. Furthermore, a Serial Bidirectional Cross-Attention (SBCA) mechanism was designed, and local queries were introduced to perform a three-stage serial interaction with multi-scale features, so that more stable local discriminative information was gradually extracted. Simultaneously, a multi-branch joint supervision strategy was employed to collaboratively optimize global and local features, so that the discriminative power and robustness of the feature representation were improved. Experimental results show that on the VeRi-776 dataset, the proposed method achieves mean Average Precision (mAP) of 84.2% and Rank-1 matching accuracy (Rank-1) of 97.7%, which are 2.8 and 0.7 percentage points higher than those of Patch-Mixup-ViT, respectively. On the Test800, Test1600 and Test2400 subsets of the VehicleID dataset, the Rank-1 of the proposed method reaches 83.8%, 80.2% and 77.2%, respectively, which are 0.8, 0.8 and 1.1 percentage points higher than those of Id-BGFNet, respectively. The results validate the effectiveness and generalization ability of the proposed method in multi-view vehicle re-identification tasks.

Key words: vehicle re-identification, Vision Transformer (ViT), multi-scale feature fusion, cross-attention mechanism, multi-branch joint supervision

摘要: 针对多视角条件下车辆重识别(Vehicle Re-ID)中间层空间信息利用不足、局部判别特征建模不充分的问题,提出一种多尺度特征引导自适应局部查询(MSAL)的车辆重识别方法。该方法以视觉Transformer(ViT)为骨干网络,从不同深度提取图像块特征,并通过特征金字塔网络的多尺度融合模块(FPN-MSF)构建层次化空间表示,以增强中间层空间信息的利用;进一步设计串行双向交叉注意力(SBCA)机制,引入局部查询与多尺度特征进行三阶段串行交互,逐步提取更稳定的局部判别信息;同时采用多分支联合监督策略,对全局与局部特征进行协同优化,以提升特征表示的判别能力与鲁棒性。实验结果表明,在VeRi-776数据集上,所提方法的平均精度均值(mAP)和首位匹配准确率(Rank-1)分别达到84.2%和97.7%,与Patch-Mixup-ViT相比,分别提高了2.8和0.7个百分点;在VehicleID数据集的Test800、Test1600和Test2400子集上,所提方法的Rank-1分别达到83.8%、80.2%和77.2%,与Id-BGFNet相比,分别提高了0.8、0.8和1.1个百分点。实验结果验证了所提方法在多视角车辆重识别任务中的有效性和泛化能力。

关键词: 车辆重识别, 视觉Transformer, 多尺度特征融合, 交叉注意力机制, 多分支联合监督

CLC Number: