Journal of Computer Applications

    Next Articles

Person re-identification method based on modality alignment and feature enhancement

  

  • Received:2026-04-15 Revised:2026-07-14 Accepted:2026-07-15 Online:2026-07-22 Published:2026-07-22

基于模态对齐与特征增强的行人重识别方法

许成杰,毕艺瀚,王铭杰,李冲,王蓉   

  1. 中国人民公安大学 信息网络安全学院,北京 100038
  • 通讯作者: 王蓉
  • 基金资助:
    国家自然科学基金;中央高校基本科研业务费专项资金项目

Abstract: To address the difficulty of learning discriminative features caused by the substantial modality discrepancy between visible and infrared images, a visible-infrared person re-identification method based on modality alignment and feature enhancement was proposed. First, a Meta-learning-guided Modality-adaptive Alignment module (M²A) was developed to capture and quantify batch-wise distribution shifts. Grayscale weights and spectral perturbation parameters were dynamically generated to guide visible images toward infrared images in the color, spatial, and spectral domains, thereby producing modality-consistent augmented images, enriching training samples, and alleviating the cross-modality discrepancy. Second, a Dynamic Feature Enhancement Module (DFEM) was constructed. Guided by modality labels, independent channel-mapping spaces were established by parallel Multi-Layer Perceptron (MLP) branches, and globally shared spatial attention was incorporated to extract modality-consistent pedestrian features, achieving dynamic enhancement in both channel and spatial dimensions. Third, a discriminative feature mining module was designed by introducing a discriminative contrastive loss into the intermediate layers of the backbone network and jointly optimizing it with classification and metric learning losses, thereby strengthening the extraction of fine-grained invariant representations. Finally, a multi-task dynamic optimization strategy was developed by combining adaptive weighting of multiple losses with a stage-wise training scheme, so that the model was progressively guided from coarse-grained global alignment to fine-grained identity mining, improving its stability and generalization capability. Under the all-search single-shot setting on the SYSU-MM01 dataset, the proposed method achieves 85.20% Rank-1 accuracy and 82.92% mean Average Precision (mAP), outperforming Multi-level Cross-Modality Joint Alignment (MCJA) by 10.72 and 11.58 percentage points, respectively. Under the visible-to-infrared setting on the RegDB dataset, it achieves 94.85% Rank-1 accuracy and 95.14% mAP, with the latter improving by 5.19 percentage points over Low-light Invariant Representation Learning (IRL). Experimental results demonstrate the effectiveness of the proposed method.

Key words: cross-modality, person re-identification, meta-learning, feature enhancement, supervised learning

摘要: 针对可见光图像和红外图像模态差异大导致判别性特征提取困难的问题,提出一种基于模态对齐与特征增强的行人重识别方法。首先,提出元学习引导的模态自适应对齐模块(M²A),结合元学习捕获并量化批次内的分布偏移,动态生成灰度权重与频谱扰动参数,引导可见光图像在色彩、空间及频谱维度上向红外图像的自适应拟合,生成具有模态一致性的增强图像,扩充训练样本并缓解模态差异。其次,构建特征动态增强模块(DFEM),利用模态标签引导并行分路多层感知机(MLP)构建独立的通道映射空间,协同全局共享空间注意力,提取模态一致的行人特征,实现特征在通道与空间维度的动态增强。再次,设计判别性特征挖掘模块,在骨干网络中层引入判别性对比损失,联合分类及度量学习损失,强化模型对细粒度不变表征的提取能力。最后,提出多任务动态优化策略,通过多任务损失的自适应加权,结合分阶段训练策略,引导模型从粗粒度全局对齐渐进转向细粒度身份挖掘,增强模型的稳定性和泛化性。在SYSU-MM01数据集的全场景单样本检索设置下,所提方法的首位命中率Rank-1和平均精度均值(mAP)分别达到85.20%和82.92%,与多层级跨模态联合对齐方法(MCJA)相比分别提升10.72个百分点和11.58个百分点;在RegDB数据集的可见光到红外检索设置下,Rank-1和mAP分别达到94.85%和95.14%,其中mAP相较于近期方法低照度不变表征学习方法(IRL)提升5.19个百分点。实验结果验证了所提方法的有效性。

关键词: 跨模态, 行人重识别, 元学习, 特征增强, 全监督学习

CLC Number: