《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (9): 2732-2740.DOI: 10.11772/j.issn.1001-9081.2025081055
• 人工智能 • 上一篇
收稿日期:2025-09-11
修回日期:2025-11-17
接受日期:2025-11-20
发布日期:2025-12-01
出版日期:2026-09-10
通讯作者:
蒋林
作者简介:张一心(1999—),女,河南周口人,硕士研究生,CCF会员,主要研究方向:神经网络模型压缩、可重构计算基金资助:
Yixin ZHANG1, Lin JIANG2(
), Yuancheng LI3, Chen JI3
Received:2025-09-11
Revised:2025-11-17
Accepted:2025-11-20
Online:2025-12-01
Published:2026-09-10
Contact:
Lin JIANG
About author:ZHANG Yixin, born in 1999, M. S. candidate. Her research interests include neural network model compression, reconfigurable computing.Supported by:摘要:
针对卷积神经网络(CNN)参数量大导致访存开销大、计算冗余和高效部署受限等问题,提出一种面向可重构结构的CNN剪枝与量化压缩方法,结合网络结构特性与硬件部署需求,从剪枝与量化两个维度协同优化。首先,提出基于特征相似的卷积层剪枝策略,依次经过特征信息评估、聚类分组、相似度计算和冗余筛选,筛选低贡献及冗余滤波器;其次,在全连接层采用渐进式阈值剪枝压缩冗余权重;再次,在量化部分利用Hessian迹构建层敏感度指标,在位宽预算下自适应地分配各层精度;最后,结合可重构结构特性,提出优化部署方案。实验结果表明,在CIFAR-10数据集上,本文方法对VGG16实现了16.20倍的压缩比,相较于APQ(Automated deep neural network Pruning and Quantization framework)的13.90倍具备更高压缩比;与16 bit固定精度的模型相比,采用本文部署策略后,剪枝后的VGG16在自重构自演化人工智能(AI)芯片上的推理时延由23.3 ms降至9.1 ms,加速比达2.56倍。本文方法在保证分类准确率的同时降低了存储与传输开销,提升了边缘设备部署效率与计算性能。
中图分类号:
张一心, 蒋林, 李远成, 纪辰. 面向可重构结构的CNN剪枝与量化压缩方法[J]. 计算机应用, 2026, 46(9): 2732-2740.
Yixin ZHANG, Lin JIANG, Yuancheng LI, Chen JI. CNN pruning and quantization compression method for reconfigurable structures[J]. Journal of Computer Applications, 2026, 46(9): 2732-2740.
| 项目 | 标签说明 | 版本型号描述 |
|---|---|---|
| CPU | 型号 | Xeon Gold 6430 |
| 物理核心数 | 16核 | |
| 内存频率 | DDR5 4 800 MHz | |
| GPU | 型号 | Nvidia RTX4090 |
| 核心数 | 16384并行运算处理核心 | |
| 显存容量 | 24 GB | |
| 系统环境 | Ubuntu22.04 | — |
| 深度学习框架 | PyTorch1.8.1 | — |
表1 实验环境配置
Tab. 1 Experimental environment configuration
| 项目 | 标签说明 | 版本型号描述 |
|---|---|---|
| CPU | 型号 | Xeon Gold 6430 |
| 物理核心数 | 16核 | |
| 内存频率 | DDR5 4 800 MHz | |
| GPU | 型号 | Nvidia RTX4090 |
| 核心数 | 16384并行运算处理核心 | |
| 显存容量 | 24 GB | |
| 系统环境 | Ubuntu22.04 | — |
| 深度学习框架 | PyTorch1.8.1 | — |
| 模型 | 方法 | Top-1准确率/% | 剪枝率/% | FLOPs/109 |
|---|---|---|---|---|
| VGG16 | 基准方法 | 93.95 | — | 0.314 |
| 本文方法 | 93.31 | 80 | 0.116 | |
| ResNet18 | 基准方法 | 93.47 | — | 0.558 |
| 本文方法 | 93.07 | 70 | 0.156 |
表2 CIFAR-10数据集上的实验结果
Tab. 2 Experimental results on CIFAR-10 dataset
| 模型 | 方法 | Top-1准确率/% | 剪枝率/% | FLOPs/109 |
|---|---|---|---|---|
| VGG16 | 基准方法 | 93.95 | — | 0.314 |
| 本文方法 | 93.31 | 80 | 0.116 | |
| ResNet18 | 基准方法 | 93.47 | — | 0.558 |
| 本文方法 | 93.07 | 70 | 0.156 |
| 网络 | 位宽 | Top-1准确率/% |
|---|---|---|
| VGG16-Pruned | 32.0 | 93.31 |
| VGG16 | 5.9 | 92.62 |
| ResNet18-Pruned | 32.0 | 93.07 |
| ResNet18 | 5.8 | 92.49 |
表3 压缩后的网络与原网络的Top-1准确率对比
Tab. 3 Comparison of Top-1 accuracy between compressed and original networks
| 网络 | 位宽 | Top-1准确率/% |
|---|---|---|
| VGG16-Pruned | 32.0 | 93.31 |
| VGG16 | 5.9 | 92.62 |
| ResNet18-Pruned | 32.0 | 93.07 |
| ResNet18 | 5.8 | 92.49 |
| 网络模型 | 基准方法 | 本文方法 | |||
|---|---|---|---|---|---|
Top-1 准确率/% | 内存占用/MB | Top-1 准确率/% | 内存占用/MB | 压缩比 | |
| VGG16 | 93.95 | 138.00 | 92.62 | 8.51 | 16.20 |
| ResNet18 | 93.47 | 44.72 | 92.49 | 5.33 | 8.38 |
表4 CNN模型联合压缩方法的效果分析
Tab. 4 Effect analysis of joint compression method of CNN models
| 网络模型 | 基准方法 | 本文方法 | |||
|---|---|---|---|---|---|
Top-1 准确率/% | 内存占用/MB | Top-1 准确率/% | 内存占用/MB | 压缩比 | |
| VGG16 | 93.95 | 138.00 | 92.62 | 8.51 | 16.20 |
| ResNet18 | 93.47 | 44.72 | 92.49 | 5.33 | 8.38 |
| 方法 | 类型 | Top-1准确率/% | 压缩后准确率 损失百分点 | 压缩比 | |
|---|---|---|---|---|---|
| 基准 | 压缩后 | ||||
| 文献[ | 混合精度量化 | 86.09 | 85.30 | 0.79↓ | 6.50 |
| 文献[ | 结构化剪枝 | 93.59 | 93.45 | 0.14↓ | 8.90 |
| 文献[ | 结构化剪枝+混合精度量化 | 93.97 | 91.26 | 2.71↓ | 10.19 |
| APQ[ | 结构化剪枝+混合精度量化 | 81.61 | 81.20 | 0.41↓ | 13.90 |
| 本文方法 | 结构化剪枝+自适应混合精度量化 | 93.95 | 92.62 | 1.33↓ | 16.20 |
表5 VGG16模型在不同压缩方法下的性能对比
Tab. 5 Performance comparison of VGG16 model under different compression methods
| 方法 | 类型 | Top-1准确率/% | 压缩后准确率 损失百分点 | 压缩比 | |
|---|---|---|---|---|---|
| 基准 | 压缩后 | ||||
| 文献[ | 混合精度量化 | 86.09 | 85.30 | 0.79↓ | 6.50 |
| 文献[ | 结构化剪枝 | 93.59 | 93.45 | 0.14↓ | 8.90 |
| 文献[ | 结构化剪枝+混合精度量化 | 93.97 | 91.26 | 2.71↓ | 10.19 |
| APQ[ | 结构化剪枝+混合精度量化 | 81.61 | 81.20 | 0.41↓ | 13.90 |
| 本文方法 | 结构化剪枝+自适应混合精度量化 | 93.95 | 92.62 | 1.33↓ | 16.20 |
| 方法 | 推理时延/ms | 加速比 |
|---|---|---|
| 未采用复用策略+统一16 bit精度 | 23.3 | 1.00 |
| 采用复用策略+统一16 bit精度 | 15.2 | 1.53 |
| 采用复用策略+混合精度量化 | 9.1 | 2.56 |
表6 不同方法剪枝后VGG16网络的推理时延对比
Tab. 6 Comparison of VGG16 network inference latency after pruning by different methods
| 方法 | 推理时延/ms | 加速比 |
|---|---|---|
| 未采用复用策略+统一16 bit精度 | 23.3 | 1.00 |
| 采用复用策略+统一16 bit精度 | 15.2 | 1.53 |
| 采用复用策略+混合精度量化 | 9.1 | 2.56 |
| [1] | Zhao X, Wang L, Zhang Y, et al. A review of convolutional neural networks in computer vision [J]. Artificial Intelligence Review, 2024, 57(4): No.99. |
| [2] | Maurício J, Domingues I, Bernardino J. Comparing vision transformers and convolutional neural networks for image classification: a literature review [J]. Applied Sciences, 2023, 13(9): No.5521. |
| [3] | Yang C B, Liu H Y. Stable low-rank cp decomposition for compression of convolutional neural networks based on sensitivity[J]. Applied Sciences, 2024, 14(4): No.1491. |
| [4] | 徐杰,郭立君,冯海,等. 轻量化的多尺度特征校准小目标检测网络[J]. 计算机应用, 2025, 45(): 228-234. |
| Xu Jie, Guo Lijun, Feng Hai, et al. Lightweight multi-scale feature calibrated small object detection network [J]. Journal of Computer Applications, 2025, 45(): 228-234. | |
| [5] | He Y, Xiao L. Structured pruning for deep convolutional neural networks: a survey [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(5): 2900-2919. |
| [6] | Liu W, Zhang M, Shi C, et al. Deep convolutional neural network compression method: tensor ring decomposition with variational Bayesian approach [J]. Neural Processing Letters, 2024, 56(2): No.103. |
| [7] | Nakata K, Miyashita D, Deguchi J, et al. Accelerating CNN inference with an adaptive quantization method using computational complexity-aware regularization [J]. IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, 2025, E108-A(2): 149-159. |
| [8] | Liu D, Zhu Y, Liu Z, et al. A survey of model compression techniques: past, present, and future [J]. Frontiers in Robotics and AI, 2025, 12: No.1518965. |
| [9] | Ahmed T, Jannat S, Rahat A, et al. Knowledge distillation and weight pruning for two-step compression of ConvNets in rice leaf disease classification [C]// NSysS 2024. New York: ACM, 2024: 72-78. |
| [10] | Heidorn C, Sabih M, Meyerhöfer N, et al. Hardware-aware evolutionary explainable filter pruning for convolutional neural networks [J]. International Journal of Parallel Programming, 2024, 52(1/2): 40-58. |
| [11] | Plochaet J, Goedemé T. Hardware-aware pruning for FPGA deep learning accelerators [C]// CVPRW 2023. Piscataway: IEEE, 2023: 4482-4490. |
| [12] | Dantas P V, Sabino da Silva W, Cordeiro L C, et al. A comprehensive review of model compression techniques in machine learning [J]. Applied Intelligence, 2024, 54(22): 11804-11844. |
| [13] | Shi Y, Bai S, Wei X, et al. Lossy and lossless (L2) post-training model size compression [C]// ICCV 2023. Piscataway: IEEE, 2023: 17500-17510. |
| [14] | 王鹏,张嘉诚,范毓洋. 适应于硬件部署的神经网络剪枝量化算法[J]. 计算机工程与科学, 2024, 46(9): 1547-1553. |
| Wang Peng, Zhang Jiacheng, Fan Yuyang. A neural network pruning and quantization algorithm for hardware deployment [J]. Computer Engineering and Science, 2024, 46(9): 1547-1553. | |
| [15] | Yang K, Jiang L, Deng J, et al. ASFRM: an array state self-feedback-based self-reconfiguration mechanism in a reconfigurable array processor [J]. Journal of Circuits, Systems and Computers, 2025, 34(14): No.2550305. |
| [16] | Hu Z, Shi Z, Karniadakis G E, et al. Hutchinson trace estimation for high-dimensional and high-order physics-informed neural networks [J]. Computer Methods in Applied Mechanics and Engineering, 2024, 424: No.116883. |
| [17] | Yang L, Zheng C, Shen X, et al. OfpCNN: on-demand fine-grained partitioning for CNN inference acceleration in heterogeneous devices [J]. IEEE Transactions on Parallel and Distributed Systems, 2023, 34(12): 3090-3103. |
| [18] | Darbani P, Beitollahi H, Lotfi-Kamran P. Rei: a reconfigurable interconnection unit for array-based CNN accelerators [J]. IEEE Transactions on Emerging Topics in Computing, 2023, 11(4): 895-906. |
| [19] | Shinde T, Bhardwaj S. Mixed-precision is all you need for efficient document image classification [C]// WACVW 2025. Piscataway: IEEE, 2025: 1195-1203. |
| [20] | Zhang D, Wang X, Wu Z, et al. A multi-scale automatic progressive pruning algorithm based on deep neural network [C]// CCC 2024. Piscataway: IEEE, 2024: 8892-8897. |
| [21] | Bai S, Chen J, Shen X, et al. Unified data-free compression: pruning and quantization without fine-tuning [C]// ICCV 2023. Piscataway: IEEE, 2023: 5853-5862. |
| [22] | Yang S, He S, Duan H, et al. APQ: automated DNN pruning and quantization for ReRAM-based accelerators [J]. IEEE Transactions on Parallel and Distributed Systems, 2023, 34(9): 2498-2511. |
| [23] | Han M, Wang L, Xiao L, et al. ReDas: a lightweight architecture for supporting fine-grained reshaping and multiple dataflows on systolic array [J]. IEEE Transactions on Computers, 2024, 73(8): 1997-2011. |
| [24] | Liu L, Jiang M, Sun J, et al. A CPU-FPGA based heterogeneous accelerator for DNA sequence alignment [C]// ICICM 2024. Piscataway: IEEE, 2024: 655-660. |
| [25] | Sun W, Liu D, Zou Z, et al. Sense: model-hardware codesign for accelerating sparse CNNs on systolic arrays [J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2023, 31(4): 470-483. |
| [26] | 魏晓辉,关泽宇,王晨洋,等. 面向脉动阵列加速器的软硬件协同容错设计[J]. 计算机科学, 2025, 52(5): 91-100. |
| Wei Xiaohui, Guan Zeyu, Wang Chenyang, et al. Hardware-software co-design fault-tolerant strategies for systolic array accelerators[J]. Computer Science, 2025, 52(5): 91-100. |
| [1] | 陈荟慧, 孙洪韬, 关柏良, 衡中青. 基于NetVLAD特征编码的古籍汉字图像检索算法[J]. 《计算机应用》唯一官方网站, 2026, 46(3): 750-757. |
| [2] | 李亚男, 郭梦阳, 邓国军, 陈允峰, 任建吉, 原永亮. 基于多模态融合特征的并分支发动机寿命预测方法[J]. 《计算机应用》唯一官方网站, 2026, 46(1): 305-313. |
| [3] | 张宏俊, 潘高军, 叶昊, 陆玉彬, 缪宜恒. 结合深度学习和张量分解的多源异构数据分析方法[J]. 《计算机应用》唯一官方网站, 2025, 45(9): 2838-2847. |
| [4] | 石超, 周昱昕, 扶倩, 唐万宇, 何凌, 李元媛. 基于骨架和3D热图的注意缺陷多动障碍患者动作识别算法[J]. 《计算机应用》唯一官方网站, 2025, 45(9): 3036-3044. |
| [5] | 彭鹏, 蔡子婷, 刘雯玲, 陈才华, 曾维, 黄宝来. 基于CNN和双向GRU混合孪生网络的语音情感识别方法[J]. 《计算机应用》唯一官方网站, 2025, 45(8): 2515-2521. |
| [6] | 林进浩, 罗川, 李天瑞, 陈红梅. 基于跨尺度注意力网络的胸部疾病分类方法[J]. 《计算机应用》唯一官方网站, 2025, 45(8): 2712-2719. |
| [7] | 张英俊, 闫薇薇, 谢斌红, 张睿, 陆望东. 梯度区分与特征范数驱动的开放世界目标检测[J]. 《计算机应用》唯一官方网站, 2025, 45(7): 2203-2210. |
| [8] | 陶永鹏, 柏诗淇, 周正文. 基于卷积和Transformer神经网络架构搜索的脑胶质瘤多组织分割网络[J]. 《计算机应用》唯一官方网站, 2025, 45(7): 2378-2386. |
| [9] | 吴宗航, 张东, 李冠宇. 基于联合自监督学习的多模态融合推荐算法[J]. 《计算机应用》唯一官方网站, 2025, 45(6): 1858-1868. |
| [10] | 龙雨菲, 牟宇辰, 刘晔. 基于张量化图卷积网络和对比学习的多源数据表示学习模型[J]. 《计算机应用》唯一官方网站, 2025, 45(5): 1372-1378. |
| [11] | 王丹, 张文豪, 彭丽娟. 基于深度学习的智能反射面辅助通信系统信道估计[J]. 《计算机应用》唯一官方网站, 2025, 45(5): 1613-1618. |
| [12] | 袁宝华, 陈佳璐, 王欢. 融合多尺度语义和双分支并行的医学图像分割网络[J]. 《计算机应用》唯一官方网站, 2025, 45(3): 988-995. |
| [13] | 耿海军, 董赟, 胡治国, 池浩田, 杨静, 尹霞. 基于Attention-1DCNN-CE的加密流量分类方法[J]. 《计算机应用》唯一官方网站, 2025, 45(3): 872-882. |
| [14] | 王地欣, 王佳昊, 李敏, 陈浩, 胡光耀, 龚宇. 面向水声通信网络的异常攻击检测[J]. 《计算机应用》唯一官方网站, 2025, 45(2): 526-533. |
| [15] | 张翰林, 王俊陆, 宋宝燕. 融合衍生特征的时间序列事件分类方法[J]. 《计算机应用》唯一官方网站, 2025, 45(2): 428-435. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||