Journal of Computer Applications

    Next Articles

Cross-domain few-shot object detection via domain-modulated hybrid prototype and feature calibration

  

  • Received:2026-04-20 Revised:2026-06-23 Accepted:2026-06-29 Online:2026-07-08 Published:2026-07-08

域调制混合原型与特征校准驱动的跨域小样本目标检测

谢斌红,屠炜明*,张睿   

  1. 太原科技大学 计算机科学与技术学院
  • 通讯作者: 屠炜明
  • 基金资助:
    山西省科技成果转化引导专项项目;吕梁市引进高层次科技人才重点研发项目;山西省产教融合研究生联合培养示范基地项目

Abstract: To address two issues in existing Cross-Domain Few-Shot Object Detection (CD-FSOD) methods, namely neglecting modality reliability differences under domain shift in multimodal prototype fusion and the susceptibility of Region of Interest (RoI) features to feature deviation and background confusion, a detection framework driven by domain-modulated hybrid prototypes and feature calibration, termed Domain-Modulated Hybrid-Prototype Calibration ViTO (DHC-ViTO), was proposed. First, a Domain-Modulated Hybrid Prototype Generation module (DM-HPG) was designed to construct visual prototypes encoding target-domain information and enhanced textual prototypes incorporating domain-related cues. A dynamic gating network was then used to adaptively adjust fusion weights according to the representation states of the two modalities in the current target domain, yielding domain-modulated hybrid prototypes with cross-domain robustness and fine-grained discriminability. Second, a Semantic Anchor-based Feature Calibration module (SAFC) was designed to use the hybrid prototypes as semantic anchors for guidance feature generation and RoI feature calibration, driving shifted RoI features toward the semantic centers of the corresponding categories. An Orthogonal Noise Decoupling (OND) strategy was further introduced to suppress background interference and improve RoI discriminability. Experimental results on six cross-domain benchmark datasets show that under the 1-shot, 5-shot, and 10-shot settings, the average mAP values of DHC-ViTO on the six target-domain datasets are 14.6%, 27.3%, and 30.8%, respectively. Compared with Cross-Domain Vision Transformer Object detector(CD-ViTO), DHC-ViTO improves the average mAP by 0.7, 1.2, and 1.2 percentage points. These results verify the effectiveness of the proposed method.

Key words: Cross-Domain Few-Shot Object Detection (CD-FSOD), domain shift, prototype learning, feature calibration, background interference

摘要: 针对现有跨域小样本目标检测方法(CD-FSOD)中,多模态原型融合未考虑域偏移下的模态可靠性差异,以及候选框(RoI)特征易产生特征偏移与背景混淆的问题,提出域调制混合原型与特征校准驱动的检测框架(DHC-ViTO)。首先,设计域调制混合原型生成模块(DM-HPG),该模块构建编码目标域信息的视觉原型,以及融入域相关线索的增强文本原型;其次,利用动态门控网络根据两种模态在当前目标域下的表征状态,自适应调节融合权重,最终生成兼具跨域鲁棒性与细粒度判别力的域调制混合原型;再次,设计基于语义锚点的特征校准模块(SAFC),针对发生偏移的RoI特征,以前述混合原型为语义锚点生成引导特征;随后,结合门控机制校准RoI特征,使它向对应类别的语义中心收敛;最后,引入正交噪声解耦(OND)策略抑制背景干扰,进一步提升RoI特征的判别性。实验结果表明,在6个跨域基准数据集上,DHC-ViTO在1-shot、5-shot和10-shot设置下的6个目标域数据集平均mAP分别达到14.6%、27.3%和30.8%;与现有先进方法CD-ViTO(Cross-Domain Vision Transformer Object detector)相比,分别提升0.7、1.2和1.2个百分点,验证了所提方法的有效性。


关键词: 跨域小样本目标检测, 域偏移, 原型学习, 特征校准, 背景干扰

CLC Number: