Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (7): 2288-2296.DOI: 10.11772/j.issn.1001-9081.2025060699

• Multimedia computing and computer simulation • Previous Articles    

Video snapshot compressive imaging reconstruction method based on dense spatio-temporal deformable attention

Xiuli DU1,2, Xing GAO1,2, Xiaoyu ZHANG1,2, Chengsheng PAN3(), Qijie ZOU2   

  1. 1.Key Laboratory of Communication and Network,Dalian University,Dalian Liaoning 116622,China
    2.School of Information Engineering,Dalian University,Dalian Liaoning 116622,China
    3.School of Electronics and Information Engineering,Nanjing University of Information Science and Technology,Nanjing Jiangsu 210044,China
  • Received:2025-06-24 Revised:2025-09-28 Accepted:2025-09-30 Online:2025-10-23 Published:2026-07-10
  • Contact: Chengsheng PAN
  • About author:DU Xiuli, born in 1977, Ph. D., professor. Her research interests include compressed sensing, perception and recognition of electroencephalogram signals.
    GAO Xing, born in 1999, M. S. candidate. His research interests include compressed sensing.
    ZHANG Xiaoyu, born in 1996, M. S. candidate. Her research interests include compressed sensing.
    ZOU Qijie, born in 1978, Ph. D., associate professor. Her research interests include computer vision, machine learning.
  • Supported by:
    Educational Department of Liaoning Province(JYTMS20230377)

基于密集时空可变形注意力的视频快照压缩成像重建方法

杜秀丽1,2, 高星1,2, 张校毓1,2, 潘成胜3(), 邹启杰2   

  1. 1.大连大学 通信与网络重点实验室,辽宁 大连 116622
    2.大连大学 信息工程学院,辽宁 大连 116622
    3.南京信息工程大学 电子与信息工程学院,南京 210044
  • 通讯作者: 潘成胜
  • 作者简介:杜秀丽(1977—),女,辽宁锦州人,教授,博士,CCF会员,主要研究方向:压缩感知、脑电信号感知与识别
    高星(1999—),男,陕西西安人,硕士研究生,主要研究方向:压缩感知
    张校毓(1996—),女,辽宁铁岭人,硕士研究生, 主要研究方向:压缩感知
    邹启杰(1978—),女,黑龙江牡丹江人,副教授,博士,主要研究方向:计算机视觉、机器学习。
  • 基金资助:
    辽宁省教育厅项目(JYTMS20230377)

Abstract:

Deep learning-based reconstruction methods for Video Snapshot Compressive Imaging (VSCI) have achieved promising results in many tasks. However, challenges such as insufficient detail recovery and high computational overhead remain in dynamic scene reconstruction. To address these issues, a VSCI reconstruction method based on dense spatio-temporal deformable attention was proposed. First, the compressed measurements and masks were input to obtain initial feature representations. Second, a deformable attention module was designed to enhance the model's ability to capture local deformations and global temporal dependencies effectively. Finally, the dense connectivity was improved by designing a channel splitting factor to enable dynamic group-wise progressive feature extraction and fusion, thereby improving feature representation ability and reducing reconstruction time. Experimental results on multiple simulated grayscale video benchmark datasets (e.g., Kobe, Runner, Drop) and color video benchmark datasets (e.g., Beauty, Bosphorus, Jockey) showed that compared with the suboptimal methods M2BA-SCI and SCT-SCI, the proposed method improves the average Peak Signal-to-Noise Ratio (PSNR) by 0.45 dB and 0.79 dB, with reconstruction times of 0.43 s and 3.37 s respectively, demonstrating that it significantly enhances the reconstruction quality and computational efficiency for complex motion scenes.

Key words: Video Snapshot Compressive Imaging (VSCI), compressed sensing, deep learning, Deformable Convolutional Network (DCN), self-attention mechanism

摘要:

基于深度学习的视频快照压缩成像(VSCI)重建方法在多数任务中已取得良好效果,但仍存在动态场景重建中细节恢复不足及计算开销大等问题。针对这些问题,提出一种基于密集时空可变形注意力的VSCI重建方法。首先,输入压缩测量值与掩码以获取初始特征信息;其次,设计一个可变形注意力模块,以有效增强模型对局部形变与全局时序依赖的捕捉能力;最后,改进密集连接,即设计通道分割因子实现动态分组递进的特征提取与融合,从而提升特征表示能力并减少重建时间。在多个模拟灰度视频基准数据集(Kobe、Runner、Drop 等)和彩色视频基准数据集(Beauty、Bosphorus、Jockey 等)上的实验结果表明,与对比方法中的次优方法M2BA?SCI和SCT-SCI相比,本文方法的平均峰值信噪比(PSNR)分别提升了0.45 dB和0.79 dB,重建时间为0.43 s和3.37 s,表明本文方法显著提升了复杂运动场景下的重建质量与计算效率。

关键词: 视频快照压缩成像, 压缩感知, 深度学习, 可变形卷积网络, 自注意力机制

CLC Number: