Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2732-2740.DOI: 10.11772/j.issn.1001-9081.2025081055
• Artificial intelligence • Previous Articles
Yixin ZHANG1, Lin JIANG2(
), Yuancheng LI3, Chen JI3
Received:2025-09-11
Revised:2025-11-17
Accepted:2025-11-20
Online:2025-12-01
Published:2026-09-10
Contact:
Lin JIANG
About author:ZHANG Yixin, born in 1999, M. S. candidate. Her research interests include neural network model compression, reconfigurable computing.Supported by:通讯作者:
蒋林
作者简介:张一心(1999—),女,河南周口人,硕士研究生,CCF会员,主要研究方向:神经网络模型压缩、可重构计算基金资助:CLC Number:
Yixin ZHANG, Lin JIANG, Yuancheng LI, Chen JI. CNN pruning and quantization compression method for reconfigurable structures[J]. Journal of Computer Applications, 2026, 46(9): 2732-2740.
张一心, 蒋林, 李远成, 纪辰. 面向可重构结构的CNN剪枝与量化压缩方法[J]. 《计算机应用》唯一官方网站, 2026, 46(9): 2732-2740.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025081055
| 项目 | 标签说明 | 版本型号描述 |
|---|---|---|
| CPU | 型号 | Xeon Gold 6430 |
| 物理核心数 | 16核 | |
| 内存频率 | DDR5 4 800 MHz | |
| GPU | 型号 | Nvidia RTX4090 |
| 核心数 | 16384并行运算处理核心 | |
| 显存容量 | 24 GB | |
| 系统环境 | Ubuntu22.04 | — |
| 深度学习框架 | PyTorch1.8.1 | — |
Tab. 1 Experimental environment configuration
| 项目 | 标签说明 | 版本型号描述 |
|---|---|---|
| CPU | 型号 | Xeon Gold 6430 |
| 物理核心数 | 16核 | |
| 内存频率 | DDR5 4 800 MHz | |
| GPU | 型号 | Nvidia RTX4090 |
| 核心数 | 16384并行运算处理核心 | |
| 显存容量 | 24 GB | |
| 系统环境 | Ubuntu22.04 | — |
| 深度学习框架 | PyTorch1.8.1 | — |
| 模型 | 方法 | Top-1准确率/% | 剪枝率/% | FLOPs/109 |
|---|---|---|---|---|
| VGG16 | 基准方法 | 93.95 | — | 0.314 |
| 本文方法 | 93.31 | 80 | 0.116 | |
| ResNet18 | 基准方法 | 93.47 | — | 0.558 |
| 本文方法 | 93.07 | 70 | 0.156 |
Tab. 2 Experimental results on CIFAR-10 dataset
| 模型 | 方法 | Top-1准确率/% | 剪枝率/% | FLOPs/109 |
|---|---|---|---|---|
| VGG16 | 基准方法 | 93.95 | — | 0.314 |
| 本文方法 | 93.31 | 80 | 0.116 | |
| ResNet18 | 基准方法 | 93.47 | — | 0.558 |
| 本文方法 | 93.07 | 70 | 0.156 |
| 网络 | 位宽 | Top-1准确率/% |
|---|---|---|
| VGG16-Pruned | 32.0 | 93.31 |
| VGG16 | 5.9 | 92.62 |
| ResNet18-Pruned | 32.0 | 93.07 |
| ResNet18 | 5.8 | 92.49 |
Tab. 3 Comparison of Top-1 accuracy between compressed and original networks
| 网络 | 位宽 | Top-1准确率/% |
|---|---|---|
| VGG16-Pruned | 32.0 | 93.31 |
| VGG16 | 5.9 | 92.62 |
| ResNet18-Pruned | 32.0 | 93.07 |
| ResNet18 | 5.8 | 92.49 |
| 网络模型 | 基准方法 | 本文方法 | |||
|---|---|---|---|---|---|
Top-1 准确率/% | 内存占用/MB | Top-1 准确率/% | 内存占用/MB | 压缩比 | |
| VGG16 | 93.95 | 138.00 | 92.62 | 8.51 | 16.20 |
| ResNet18 | 93.47 | 44.72 | 92.49 | 5.33 | 8.38 |
Tab. 4 Effect analysis of joint compression method of CNN models
| 网络模型 | 基准方法 | 本文方法 | |||
|---|---|---|---|---|---|
Top-1 准确率/% | 内存占用/MB | Top-1 准确率/% | 内存占用/MB | 压缩比 | |
| VGG16 | 93.95 | 138.00 | 92.62 | 8.51 | 16.20 |
| ResNet18 | 93.47 | 44.72 | 92.49 | 5.33 | 8.38 |
| 方法 | 类型 | Top-1准确率/% | 压缩后准确率 损失百分点 | 压缩比 | |
|---|---|---|---|---|---|
| 基准 | 压缩后 | ||||
| 文献[ | 混合精度量化 | 86.09 | 85.30 | 0.79↓ | 6.50 |
| 文献[ | 结构化剪枝 | 93.59 | 93.45 | 0.14↓ | 8.90 |
| 文献[ | 结构化剪枝+混合精度量化 | 93.97 | 91.26 | 2.71↓ | 10.19 |
| APQ[ | 结构化剪枝+混合精度量化 | 81.61 | 81.20 | 0.41↓ | 13.90 |
| 本文方法 | 结构化剪枝+自适应混合精度量化 | 93.95 | 92.62 | 1.33↓ | 16.20 |
Tab. 5 Performance comparison of VGG16 model under different compression methods
| 方法 | 类型 | Top-1准确率/% | 压缩后准确率 损失百分点 | 压缩比 | |
|---|---|---|---|---|---|
| 基准 | 压缩后 | ||||
| 文献[ | 混合精度量化 | 86.09 | 85.30 | 0.79↓ | 6.50 |
| 文献[ | 结构化剪枝 | 93.59 | 93.45 | 0.14↓ | 8.90 |
| 文献[ | 结构化剪枝+混合精度量化 | 93.97 | 91.26 | 2.71↓ | 10.19 |
| APQ[ | 结构化剪枝+混合精度量化 | 81.61 | 81.20 | 0.41↓ | 13.90 |
| 本文方法 | 结构化剪枝+自适应混合精度量化 | 93.95 | 92.62 | 1.33↓ | 16.20 |
| 方法 | 推理时延/ms | 加速比 |
|---|---|---|
| 未采用复用策略+统一16 bit精度 | 23.3 | 1.00 |
| 采用复用策略+统一16 bit精度 | 15.2 | 1.53 |
| 采用复用策略+混合精度量化 | 9.1 | 2.56 |
Tab. 6 Comparison of VGG16 network inference latency after pruning by different methods
| 方法 | 推理时延/ms | 加速比 |
|---|---|---|
| 未采用复用策略+统一16 bit精度 | 23.3 | 1.00 |
| 采用复用策略+统一16 bit精度 | 15.2 | 1.53 |
| 采用复用策略+混合精度量化 | 9.1 | 2.56 |
| [1] | Zhao X, Wang L, Zhang Y, et al. A review of convolutional neural networks in computer vision [J]. Artificial Intelligence Review, 2024, 57(4): No.99. |
| [2] | Maurício J, Domingues I, Bernardino J. Comparing vision transformers and convolutional neural networks for image classification: a literature review [J]. Applied Sciences, 2023, 13(9): No.5521. |
| [3] | Yang C B, Liu H Y. Stable low-rank cp decomposition for compression of convolutional neural networks based on sensitivity[J]. Applied Sciences, 2024, 14(4): No.1491. |
| [4] | 徐杰,郭立君,冯海,等. 轻量化的多尺度特征校准小目标检测网络[J]. 计算机应用, 2025, 45(): 228-234. |
| Xu Jie, Guo Lijun, Feng Hai, et al. Lightweight multi-scale feature calibrated small object detection network [J]. Journal of Computer Applications, 2025, 45(): 228-234. | |
| [5] | He Y, Xiao L. Structured pruning for deep convolutional neural networks: a survey [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(5): 2900-2919. |
| [6] | Liu W, Zhang M, Shi C, et al. Deep convolutional neural network compression method: tensor ring decomposition with variational Bayesian approach [J]. Neural Processing Letters, 2024, 56(2): No.103. |
| [7] | Nakata K, Miyashita D, Deguchi J, et al. Accelerating CNN inference with an adaptive quantization method using computational complexity-aware regularization [J]. IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, 2025, E108-A(2): 149-159. |
| [8] | Liu D, Zhu Y, Liu Z, et al. A survey of model compression techniques: past, present, and future [J]. Frontiers in Robotics and AI, 2025, 12: No.1518965. |
| [9] | Ahmed T, Jannat S, Rahat A, et al. Knowledge distillation and weight pruning for two-step compression of ConvNets in rice leaf disease classification [C]// NSysS 2024. New York: ACM, 2024: 72-78. |
| [10] | Heidorn C, Sabih M, Meyerhöfer N, et al. Hardware-aware evolutionary explainable filter pruning for convolutional neural networks [J]. International Journal of Parallel Programming, 2024, 52(1/2): 40-58. |
| [11] | Plochaet J, Goedemé T. Hardware-aware pruning for FPGA deep learning accelerators [C]// CVPRW 2023. Piscataway: IEEE, 2023: 4482-4490. |
| [12] | Dantas P V, Sabino da Silva W, Cordeiro L C, et al. A comprehensive review of model compression techniques in machine learning [J]. Applied Intelligence, 2024, 54(22): 11804-11844. |
| [13] | Shi Y, Bai S, Wei X, et al. Lossy and lossless (L2) post-training model size compression [C]// ICCV 2023. Piscataway: IEEE, 2023: 17500-17510. |
| [14] | 王鹏,张嘉诚,范毓洋. 适应于硬件部署的神经网络剪枝量化算法[J]. 计算机工程与科学, 2024, 46(9): 1547-1553. |
| Wang Peng, Zhang Jiacheng, Fan Yuyang. A neural network pruning and quantization algorithm for hardware deployment [J]. Computer Engineering and Science, 2024, 46(9): 1547-1553. | |
| [15] | Yang K, Jiang L, Deng J, et al. ASFRM: an array state self-feedback-based self-reconfiguration mechanism in a reconfigurable array processor [J]. Journal of Circuits, Systems and Computers, 2025, 34(14): No.2550305. |
| [16] | Hu Z, Shi Z, Karniadakis G E, et al. Hutchinson trace estimation for high-dimensional and high-order physics-informed neural networks [J]. Computer Methods in Applied Mechanics and Engineering, 2024, 424: No.116883. |
| [17] | Yang L, Zheng C, Shen X, et al. OfpCNN: on-demand fine-grained partitioning for CNN inference acceleration in heterogeneous devices [J]. IEEE Transactions on Parallel and Distributed Systems, 2023, 34(12): 3090-3103. |
| [18] | Darbani P, Beitollahi H, Lotfi-Kamran P. Rei: a reconfigurable interconnection unit for array-based CNN accelerators [J]. IEEE Transactions on Emerging Topics in Computing, 2023, 11(4): 895-906. |
| [19] | Shinde T, Bhardwaj S. Mixed-precision is all you need for efficient document image classification [C]// WACVW 2025. Piscataway: IEEE, 2025: 1195-1203. |
| [20] | Zhang D, Wang X, Wu Z, et al. A multi-scale automatic progressive pruning algorithm based on deep neural network [C]// CCC 2024. Piscataway: IEEE, 2024: 8892-8897. |
| [21] | Bai S, Chen J, Shen X, et al. Unified data-free compression: pruning and quantization without fine-tuning [C]// ICCV 2023. Piscataway: IEEE, 2023: 5853-5862. |
| [22] | Yang S, He S, Duan H, et al. APQ: automated DNN pruning and quantization for ReRAM-based accelerators [J]. IEEE Transactions on Parallel and Distributed Systems, 2023, 34(9): 2498-2511. |
| [23] | Han M, Wang L, Xiao L, et al. ReDas: a lightweight architecture for supporting fine-grained reshaping and multiple dataflows on systolic array [J]. IEEE Transactions on Computers, 2024, 73(8): 1997-2011. |
| [24] | Liu L, Jiang M, Sun J, et al. A CPU-FPGA based heterogeneous accelerator for DNA sequence alignment [C]// ICICM 2024. Piscataway: IEEE, 2024: 655-660. |
| [25] | Sun W, Liu D, Zou Z, et al. Sense: model-hardware codesign for accelerating sparse CNNs on systolic arrays [J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2023, 31(4): 470-483. |
| [26] | 魏晓辉,关泽宇,王晨洋,等. 面向脉动阵列加速器的软硬件协同容错设计[J]. 计算机科学, 2025, 52(5): 91-100. |
| Wei Xiaohui, Guan Zeyu, Wang Chenyang, et al. Hardware-software co-design fault-tolerant strategies for systolic array accelerators[J]. Computer Science, 2025, 52(5): 91-100. |
| [1] | Huihui CHEN, Hongtao SUN, Boliang GUAN, Zhongqing HENG. Chinese character image retrieval algorithm in ancient books based on NetVLAD feature encoding [J]. Journal of Computer Applications, 2026, 46(3): 750-757. |
| [2] | Yanan LI, Mengyang GUO, Guojun DENG, Yunfeng CHEN, Jianji REN, Yongliang YUAN. Method for life prediction of parallel branching engine based on multi-modal fusion features [J]. Journal of Computer Applications, 2026, 46(1): 305-313. |
| [3] | Hongjun ZHANG, Gaojun PAN, Hao YE, Yubin LU, Yiheng MIAO. Multi-source heterogeneous data analysis method combining deep learning and tensor decomposition [J]. Journal of Computer Applications, 2025, 45(9): 2838-2847. |
| [4] | Chao SHI, Yuxin ZHOU, Qian FU, Wanyu TANG, Ling HE, Yuanyuan LI. Action recognition algorithm for ADHD patients using skeleton and 3D heatmap [J]. Journal of Computer Applications, 2025, 45(9): 3036-3044. |
| [5] | Peng PENG, Ziting CAI, Wenling LIU, Caihua CHEN, Wei ZENG, Baolai HUANG. Speech emotion recognition method based on hybrid Siamese network with CNN and bidirectional GRU [J]. Journal of Computer Applications, 2025, 45(8): 2515-2521. |
| [6] | Jinhao LIN, Chuan LUO, Tianrui LI, Hongmei CHEN. Thoracic disease classification method based on cross-scale attention network [J]. Journal of Computer Applications, 2025, 45(8): 2712-2719. |
| [7] | Yongpeng TAO, Shiqi BAI, Zhengwen ZHOU. Neural architecture search for multi-tissue segmentation using convolutional and transformer-based networks in glioma segmentation [J]. Journal of Computer Applications, 2025, 45(7): 2378-2386. |
| [8] | Yingjun ZHANG, Weiwei YAN, Binhong XIE, Rui ZHANG, Wangdong LU. Gradient-discriminative and feature norm-driven open-world object detection [J]. Journal of Computer Applications, 2025, 45(7): 2203-2210. |
| [9] | Dan WANG, Wenhao ZHANG, Lijuan PENG. Channel estimation of reconfigurable intelligent surface assisted communication system based on deep learning [J]. Journal of Computer Applications, 2025, 45(5): 1613-1618. |
| [10] | Baohua YUAN, Jialu CHEN, Huan WANG. Medical image segmentation network integrating multi-scale semantics and parallel double-branch [J]. Journal of Computer Applications, 2025, 45(3): 988-995. |
| [11] | Sheng YANG, Yan LI. Contrastive knowledge distillation method for object detection [J]. Journal of Computer Applications, 2025, 45(2): 354-361. |
| [12] | Dixin WANG, Jiahao WANG, Min LI, Hao CHEN, Guangyao HU, Yu GONG. Abnormal attack detection for underwater acoustic communication network [J]. Journal of Computer Applications, 2025, 45(2): 526-533. |
| [13] | Yonghong FAN, Heming HUANG. CnnPRL: progressive representation learning method for speech emotion recognition [J]. Journal of Computer Applications, 2025, 45(12): 3804-3812. |
| [14] | Jianhua REN, Jiahui CAO, Di JIA. Hand pose estimation based on mask prompts and attention [J]. Journal of Computer Applications, 2025, 45(12): 4012-4020. |
| [15] | Xinran XU, Shaobing ZHANG, Miao CHENG, Yang ZHANG, Shang ZENG. Bearings fault diagnosis method based on multi-pathed hierarchical mixture-of-experts model [J]. Journal of Computer Applications, 2025, 45(1): 59-68. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||