Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2787-2792.DOI: 10.11772/j.issn.1001-9081.2025081009
• Artificial intelligence • Previous Articles
Received:2025-09-03
Revised:2025-11-14
Accepted:2025-11-18
Online:2025-12-01
Published:2026-09-10
Contact:
Zhangjian JI
About author:JI Zhangjian, born in 1983, Ph. D., associate professor. His research interests include computer vision, machine learning.Supported by:通讯作者:
姬张建
作者简介:姬张建(1983—),男,陕西澄城人,副教授,博士,CCF会员,主要研究方向:计算机视觉、机器学习基金资助:CLC Number:
Zhangjian JI, Siyuan WANG. Lightweight human pose estimation framework based on wavelet attention mechanism with enhanced low frequency[J]. Journal of Computer Applications, 2026, 46(9): 2787-2792.
姬张建, 王思源. 基于增强低频小波注意力机制的轻量化人体姿态估计框架[J]. 《计算机应用》唯一官方网站, 2026, 46(9): 2787-2792.
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2025081009
| 类型 | 方法 | 主干网络 | mAP/% | AP50/% | AP75/% | AP_M/% | AP_L/% | 参数量/106 | 浮点运算量/GFLOPs |
|---|---|---|---|---|---|---|---|---|---|
| 轻量级模型 | HRFormer-T*[ | HRFormer-T[ | 72.4 | 89.3 | 79.0 | 68.2 | 79.7 | 2.8 | 1.8 |
| MSPose-T[ | TokenPose-T[ | 67.1 | 87.3 | 75.3 | 64.4 | 73.0 | 5.8 | 1.3 | |
| MobileNet[ | MobileNetV2[ | 64.8 | 87.4 | 72.5 | — | — | 9.6 | 1.6 | |
| ShuffleNet[ | ShuffleNetV2[ | 60.2 | 85.7 | 67.2 | — | — | 7.6 | 1.4 | |
| Lite Pose[ | LitePose-XS[ | 40.6 | — | — | — | — | 1.7 | 1.2 | |
| LMFormer[ | LMFormer-L[ | 68.9 | 88.3 | 76.4 | — | — | 4.1 | 1.4 | |
| Lite-HRNet[ | HRNet | 64.8 | 86.7 | 73.0 | 62.1 | 70.5 | 1.1 | 0.2 | |
| TokenPose-T[ | Transformer[ | 65.6 | 86.4 | 73.0 | 63.1 | 71.5 | 5.8 | 1.3 | |
| 大尺寸模型 | SimpleBaseline+[ | ResNet-152 | 73.7 | 91.9 | 81.8 | 70.3 | 80.0 | 68.6 | 15.7 |
| HRNet[ | HRNet-W32 | 73.4 | 89.5 | 80.7 | 70.2 | 80.1 | 28.5 | 7.1 | |
| DARK+[ | HRNet-W48 | 76.2 | 92.5 | 83.6 | 72.5 | 82.4 | 63.6 | 32.9 | |
| TokenPose-B*[ | HRNet-W32 | 74.7 | 89.8 | 81.4 | 71.3 | 81.4 | 13.5 | 5.7 | |
| TokenPose-L[ | HRNet-W48 | 75.1 | 92.1 | 82.5 | 71.7 | 81.1 | 27.5 | 11.0 | |
| PoseTrans[ | ResNet-101 | 72.7 | 90.0 | 80.7 | 69.5 | 78.8 | 53.0 | 12.4 | |
| RLE[ | HRNet-W48 | 75.7 | 92.3 | 82.9 | 72.3 | 81.3 | 63.6 | 32.9 | |
| RSN[ | RSN-50 | 72.5 | 93.0 | 81.3 | 69.9 | 76.5 | 25.7 | 6.4 | |
| FCAC | HRNet-W32 | 76.9 | 93.8 | 83.9 | 73.9 | 81.2 | 34.8 | 6.6 |
Tab. 1 Comparison of different model results on COCO dataset
| 类型 | 方法 | 主干网络 | mAP/% | AP50/% | AP75/% | AP_M/% | AP_L/% | 参数量/106 | 浮点运算量/GFLOPs |
|---|---|---|---|---|---|---|---|---|---|
| 轻量级模型 | HRFormer-T*[ | HRFormer-T[ | 72.4 | 89.3 | 79.0 | 68.2 | 79.7 | 2.8 | 1.8 |
| MSPose-T[ | TokenPose-T[ | 67.1 | 87.3 | 75.3 | 64.4 | 73.0 | 5.8 | 1.3 | |
| MobileNet[ | MobileNetV2[ | 64.8 | 87.4 | 72.5 | — | — | 9.6 | 1.6 | |
| ShuffleNet[ | ShuffleNetV2[ | 60.2 | 85.7 | 67.2 | — | — | 7.6 | 1.4 | |
| Lite Pose[ | LitePose-XS[ | 40.6 | — | — | — | — | 1.7 | 1.2 | |
| LMFormer[ | LMFormer-L[ | 68.9 | 88.3 | 76.4 | — | — | 4.1 | 1.4 | |
| Lite-HRNet[ | HRNet | 64.8 | 86.7 | 73.0 | 62.1 | 70.5 | 1.1 | 0.2 | |
| TokenPose-T[ | Transformer[ | 65.6 | 86.4 | 73.0 | 63.1 | 71.5 | 5.8 | 1.3 | |
| 大尺寸模型 | SimpleBaseline+[ | ResNet-152 | 73.7 | 91.9 | 81.8 | 70.3 | 80.0 | 68.6 | 15.7 |
| HRNet[ | HRNet-W32 | 73.4 | 89.5 | 80.7 | 70.2 | 80.1 | 28.5 | 7.1 | |
| DARK+[ | HRNet-W48 | 76.2 | 92.5 | 83.6 | 72.5 | 82.4 | 63.6 | 32.9 | |
| TokenPose-B*[ | HRNet-W32 | 74.7 | 89.8 | 81.4 | 71.3 | 81.4 | 13.5 | 5.7 | |
| TokenPose-L[ | HRNet-W48 | 75.1 | 92.1 | 82.5 | 71.7 | 81.1 | 27.5 | 11.0 | |
| PoseTrans[ | ResNet-101 | 72.7 | 90.0 | 80.7 | 69.5 | 78.8 | 53.0 | 12.4 | |
| RLE[ | HRNet-W48 | 75.7 | 92.3 | 82.9 | 72.3 | 81.3 | 63.6 | 32.9 | |
| RSN[ | RSN-50 | 72.5 | 93.0 | 81.3 | 69.9 | 76.5 | 25.7 | 6.4 | |
| FCAC | HRNet-W32 | 76.9 | 93.8 | 83.9 | 73.9 | 81.2 | 34.8 | 6.6 |
| 方法 | 主干网络 | 人体不同部分的识别准确率 | 平均 准确率 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Head | Sho. | Elb. | Wri. | Hip | Knee | Ank. | |||
| HRNet[ | HRNet-W32 | 97.1 | 95.9 | 90.3 | 86.4 | 89.1 | 87.1 | 83.3 | 90.3 |
| TokenPose-T[ | HRFormer-T | 97.1 | 95.9 | 91.0 | 85.8 | 89.5 | 86.1 | 82.7 | 90.2 |
| SimpleBaseline+[ | ResNet-152 | 97.0 | 95.9 | 90.3 | 85.0 | 89.2 | 85.3 | 81.3 | 89.6 |
| DARK+[ | HRNet-W48 | 97.2 | 95.9 | 91.2 | 86.7 | 89.7 | 86.7 | 84.0 | 90.6 |
| FCAC | HRNet-W32 | 97.2 | 95.9 | 90.4 | 86.3 | 89.2 | 87.0 | 83.3 | 90.5 |
Tab. 2 Comparison of different model results on MPII dataset
| 方法 | 主干网络 | 人体不同部分的识别准确率 | 平均 准确率 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Head | Sho. | Elb. | Wri. | Hip | Knee | Ank. | |||
| HRNet[ | HRNet-W32 | 97.1 | 95.9 | 90.3 | 86.4 | 89.1 | 87.1 | 83.3 | 90.3 |
| TokenPose-T[ | HRFormer-T | 97.1 | 95.9 | 91.0 | 85.8 | 89.5 | 86.1 | 82.7 | 90.2 |
| SimpleBaseline+[ | ResNet-152 | 97.0 | 95.9 | 90.3 | 85.0 | 89.2 | 85.3 | 81.3 | 89.6 |
| DARK+[ | HRNet-W48 | 97.2 | 95.9 | 91.2 | 86.7 | 89.7 | 86.7 | 84.0 | 90.6 |
| FCAC | HRNet-W32 | 97.2 | 95.9 | 90.4 | 86.3 | 89.2 | 87.0 | 83.3 | 90.5 |
| WT | DSC | CFAF | SENet | mAP/% |
|---|---|---|---|---|
| — | — | — | — | 73.4 |
| — | | — | — | 73.8 |
| | | — | — | 76.1 |
| | | — | | 76.3 |
| | | | — | 76.6 |
| | | | | 76.9 |
Tab. 3 Ablation analysis of different FCAC modules
| WT | DSC | CFAF | SENet | mAP/% |
|---|---|---|---|---|
| — | — | — | — | 73.4 |
| — | | — | — | 73.8 |
| | | — | — | 76.1 |
| | | — | | 76.3 |
| | | | — | 76.6 |
| | | | | 76.9 |
| El | Eh | mAP/% |
|---|---|---|
| — | — | 76.1 |
| — | | 76.7 |
| | — | 76.9 |
| | | 76.8 |
Tab. 4 Impact of high- and low-frequency enhancement on FCAC
| El | Eh | mAP/% |
|---|---|---|
| — | — | 76.1 |
| — | | 76.7 |
| | — | 76.9 |
| | | 76.8 |
| 策略 | mAP/% | 参数量/106 | 浮点运算量/GFLOPs |
|---|---|---|---|
| All | 73.9 | 18.5 | 1.5 |
| Bottleneck | 73.5 | 28.9 | 6.9 |
| OnlyD | 76.9 | 34.8 | 6.6 |
Tab. 5 Impact of replacement strategies on performance
| 策略 | mAP/% | 参数量/106 | 浮点运算量/GFLOPs |
|---|---|---|---|
| All | 73.9 | 18.5 | 1.5 |
| Bottleneck | 73.5 | 28.9 | 6.9 |
| OnlyD | 76.9 | 34.8 | 6.6 |
| [1] | Li K, Wang S, Zhang X, et al. Pose recognition with cascade Transformers [C]// CVPR 2021. Piscataway: IEEE, 2021: 1944-1953. |
| [2] | Cao Z, Simo T, Wei S E, et al. Realtime multi-person 2D pose estimation using part affinity fields [C]// CVPR 2017. Piscataway: IEEE, 2017: 1302-1310. |
| [3] | Kocabas M, Karagoz S, Akbas E. MultiPoseNet: fast multi-person pose estimation using pose residual network [C]// ECCV 2018, LNCS 11215. Cham: Springer, 2018: 437-453. |
| [4] | Papandreou G, Zhu T, Chen L C, et al. PersonLab: person pose estimation and instance segmentation with a bottom-up, part-based, geometric embedding model [C]// ECCV 2018, LNCS 11218. Cham: Springer, 2018: 282-299. |
| [5] | Ahn D, Kim S, Hong H, et al. STAR-Transformer: a spatio-temporal cross attention Transformer for human action recognition[C]// WACV 2023. Piscataway: IEEE, 2023: 3319-3328. |
| [6] | Zhou T, Yang Y, Wang W. Differentiable multi-granularity human parsing [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(7): 8296-8310. |
| [7] | Xie X, Bhatnagar B L, Pons-Moll G. Visibility aware human-object interaction tracking from single RGB camera [C]// CVPR 2023. Piscataway: IEEE, 2023: 4757-4768. |
| [8] | Guo W, Du Y, Shen X, et al. Back to MLP: a simple baseline for human motion prediction [C]// WACV 2023. Piscataway: IEEE, 2023: 4798-4808. |
| [9] | Yan X, Xu Y, Chen C, et al. Privacy preserving for AI-based 3D human pose recovery and retargeting [J]. ISA Transactions, 2023, 141: 132-142. |
| [10] | Toshev A, Szegedy C. DeepPose: human pose estimation via deep neural networks [C]// CVPR 2014. Piscataway: IEEE, 2014: 1653-1660. |
| [11] | Pfister T, Simonyan K, Charles J, et al. Deep convolutional neural networks for efficient pose estimation in gesture videos [C]// ACCV 2014, LNCS 9003. Cham: Springer, 2015: 538-552. |
| [12] | Newell A, Huang Z, Deng J. Associative embedding: end-to-end learning for joint detection and grouping [C]// NeurIPS 2017. Red Hook: Curran Associates Inc., 2017: 2274-2284. |
| [13] | Insafutdinov E, Pishchulin L, Andres B, et al. DeeperCut: a deeper, stronger, and faster multi-person pose estimation model[C]// ECCV 2016, LNCS 9910. Cham: Springer, 2016: 34-50. |
| [14] | 孔英会,秦胤峰,张珂. 深度学习二维人体姿态估计方法综述[J]. 中国图象图形学报, 2023, 28(7): 1965-1989. |
| Kong Yinghui, Qin Yinfeng, Zhang Ke. Deep learning based two-dimension human pose estimation: a critical analysis[J]. Journal of Image and Graphics, 2023, 28(7): 1965-1989. | |
| [15] | Sun K, Xiao B, Liu D, et al. Deep high-resolution representation learning for human pose estimation [C]// CVPR 2019. Piscataway: IEEE, 2019: 5686-5696. |
| [16] | Howard A, Sandler M, Chen B, et al. Searching for MobileNetV3[C]// ICCV 2019. Piscataway: IEEE, 2019: 1314-1324. |
| [17] | Li Q, Shen L, Guo S, et al. Wavelet integrated CNNs for noise-robust image classification [C]// CVPR 2020. Piscataway: IEEE, 2020: 7243-7252. |
| [18] | Yu C, Xiao B, Gao C, et al. Lite-HRNet: a lightweight high-resolution network [C]// CVPR 2021. Piscataway: IEEE, 2021: 10435-10445. |
| [19] | Liu Y, Hua J. L-HRNet: a lightweight high-resolution network for human pose estimation [C]// ICIIBMS 2023. Piscataway: IEEE, 2023: 219-224. |
| [20] | Finder S E, Amoyal R, Treister E, et al. Wavelet convolutions for large receptive fields [C]// ECCV 2024, LNCS 15112. Cham: Springer, 2025: 363-380. |
| [21] | Chollet F. Xception: deep learning with depthwise separable convolutions [C]// CVPR 2017. Piscataway: IEEE, 2017: 1800-1807. |
| [22] | Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]// NeurIPS 2017. Red Hook: Curran Associates Inc., 2017: 6000-6010. |
| [23] | Tammina S. Transfer learning using VGG-16 with deep convolutional neural network for classifying images [J]. International Journal of Scientific and Research Publications, 2019, 9(10): 143-150. |
| [24] | He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C]// CVPR 2016. Piscataway: IEEE, 2016: 770-778. |
| [25] | Fang H S, Xie S, Tai Y W, et al. RMPE: regional multi-person pose estimation [C]// ICCV 2017. Piscataway: IEEE, 2017: 2353-2362. |
| [26] | Yuan Y, Fu R, Huang L, et al. HRFormer: high-resolution vision Transformer for dense predict [C]// NeurIPS 2021. Red Hook: Curran Associates Inc., 2021: 7281-7293. |
| [27] | Newell A, Yang K, Deng J. Stacked Hourglass networks for human pose estimation [C]// ECCV 2016. Cham: Springer, 2016: 483-499. |
| [28] | 盛爱兰,李舜酩. 小波分析及其应用的研究现状和发展趋势[J]. 淄博学院学报(自然科学与工程版), 2001, 3(4): 51-56. |
| Sheng Ailan, Li Shunming. The current research situation and development trend of wavelet analysis and application [J]. Journal of Zibo University (Natural Science and Engineering Edition), 2001, 3(4): 51-56. | |
| [29] | 吕印晓,刘芳. 一种基于小波多尺度边缘表示的图像滤波方法[J]. 计算机工程与应用, 2001, 37(13): 129-131. |
| Yinxiao Lyu, Liu Fang. A filtering method of images based on the multiscale edge representation [J]. Computer Engineering and Applications, 2001, 37(13): 129-131. | |
| [30] | Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module [C]// ECCV 2018, LNCS 11211. Cham: Springer, 2018: 3-19. |
| [31] | Hu J, Shen L, Sun G. Squeeze-and-excitation networks [C]// CVPR 2018. Piscataway: IEEE, 2018: 7132-7141. |
| [32] | Lin T Y, Maire M, Belongie S, et al. Microsoft COCO: common objects in context [C]// ECCV 2014, LNCS 8693. Cham: Springer, 2014: 740-755. |
| [33] | Yuan X, Cheng P, Han S. Multi-supervision Transformer combining bounding box and mask for data-limited pose estimation[J]. Neurocomputing, 2024, 571: No.127209. |
| [34] | Li Y, Zhang S, Wang Z, et al. TokenPose: learning keypoint tokens for human pose estimation [C]// CVPR 2021. Piscataway: IEEE, 2021: 11293-11302. |
| [35] | Howard A G, Zhu M, Chen B, et al. MobileNets: efficient convolutional neural networks for mobile vision applications [PP/OL]. arXiv (2017-04-17) [2025-06-21].. |
| [36] | Sandler M, Howard A, Zhu M, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]// CVPR 2018. Piscataway: IEEE, 2018: 4510-4520. |
| [37] | Zhang X, Zhou X, Lin M, et al. ShuffleNet: an extremely efficient convolutional neural network for mobile devices [C]// CVPR 2018. Piscataway: IEEE, 2018: 6848-6856. |
| [38] | Ma N, Zhang X, Zheng H T, et al. ShuffleNet V2: practical guidelines for efficient CNN architecture design [C]// ECCV 2018, LNCS 11218. Cham: Springer, 2018: 122-138. |
| [39] | Wang Y, Li M, Cai H, et al. Lite Pose: efficient architecture design for 2D human pose estimation [C]// CVPR 2022. Piscataway: IEEE, 2022: 13116-13126. |
| [40] | Li B, Tang S, Li W. LMFormer: lightweight and multi-feature perspective via Transformer for human pose estimation [J]. Neurocomputing, 2024, 594: No.127884. |
| [41] | Xiao B, Wu H, Wei Y. Simple baselines for human pose estimation and tracking [C]// ECCV 2018, LNCS 11210. Cham: Springer, 2018: 472-487. |
| [42] | Zhang F, Zhu X, Dai H, et al. Distribution-aware coordinate representation for human pose estimation [C]// CVPR 2020. Piscataway: IEEE, 2020: 7091-7100. |
| [43] | Jiang W, Jin S, Liu W, et al. PoseTrans: a simple yet effective pose transformation augmentation for human pose estimation [C]// ECCV 2022, LNCS 13665. Cham: Springer, 2022: 643-659. |
| [44] | Houston B, Nielsen M B, Batty C, et al. Hierarchical RLE level set: a compact and versatile deformable surface representation [J]. ACM Transactions on Graphics, 2006, 25(1): 151-175. |
| [45] | Sun P, Wang W, Chai Y, et al. RSN: range sparse net for efficient, accurate LiDAR 3D object detection [C]// CVPR 2021. Piscataway: IEEE, 2021: 5721-5730. |
| [1] | Jinxiao ZHANG, Chenglong LI, Xinyan GAO, Ming ZHANG. 3D human pose estimation model based on temporal-spatial feature pyramid network and multi-hypothesis interaction mechanism [J]. Journal of Computer Applications, 2026, 46(6): 1965-1972. |
| [2] | Liwan YAO, Hailong LIU, Zhangfan ZENG. Frequency-domain driven and diffusion-based fusion for sonar image enhancement algorithm [J]. Journal of Computer Applications, 2026, 46(6): 1947-1955. |
| [3] | Chao LYU, Geyao MA. Lightweight human pose estimation network based on redundant feature suppression [J]. Journal of Computer Applications, 2026, 46(6): 1973-1980. |
| [4] | Qiuyan YIN, Jing DING, Zhigang NIE. YOLO-AirPose: human pose estimation algorithm in UAV aerial view [J]. Journal of Computer Applications, 2026, 46(6): 1989-1997. |
| [5] | Yu CHEN, Shuaikang QI, Liwei XU, Haotian ZHU. Social bot detection framework fusing multi-scale wavelet enhancement and self-supervised learning [J]. Journal of Computer Applications, 2026, 46(6): 1756-1766. |
| [6] | Ying JING, Ran LI, Zhuo JIANG, Ziyang FU, Jingyi DU, Qi LIU, Jihang LIU. SAM Meibomian gland unified dense segmentation method with introduction of automatic prompt encoder [J]. Journal of Computer Applications, 2026, 46(5): 1667-1676. |
| [7] | Kaiyan CUI, Shuna WEI. Wavelet-domain sparse Bayesian learning for uncertainty-aware MRI reconstruction [J]. Journal of Computer Applications, 2026, 46(5): 1634-1646. |
| [8] | Xiang BAI, Juchuan LI, Huimin WANG, Chao JING, Jian NIU, Xingzhong ZHANG, Yongqiang CHENG. Power image retrieval method based on improved Swin Transformer [J]. Journal of Computer Applications, 2026, 46(4): 1334-1343. |
| [9] | Xiaowei LA, Lihua HU, Jianhua HU, Xiaoling YAO, Xinbo WANG. Low-overlap point cloud registration network integrating position encoding and overlap masks [J]. Journal of Computer Applications, 2026, 46(2): 536-545. |
| [10] | Yanan LI, Mengyang GUO, Guojun DENG, Yunfeng CHEN, Jianji REN, Yongliang YUAN. Method for life prediction of parallel branching engine based on multi-modal fusion features [J]. Journal of Computer Applications, 2026, 46(1): 305-313. |
| [11] | Yihan WANG, Chong LU, Zhongyuan CHEN. Multimodal sentiment analysis model with cross-modal text information enhancement [J]. Journal of Computer Applications, 2025, 45(7): 2237-2244. |
| [12] | Hailin XIAO, Xiangting KONG, Yu WANG, Di ZHOU, Xiaoming DAI. Image watermarking algorithm based on improved singular value decomposition and Haar wavelet transform [J]. Journal of Computer Applications, 2025, 45(3): 896-903. |
| [13] | Benjie SHE, Shuzhi SU, Yanmin ZHU, Jian HUA, Chao WANG. Lightweight pose estimation network based on non-globally dependent integral regression [J]. Journal of Computer Applications, 2025, 45(3): 972-977. |
| [14] | Weigang LI, Wenjie CAO, Jinling LI. Multi-stage point cloud completion network based on adaptive neighborhood feature fusion [J]. Journal of Computer Applications, 2025, 45(10): 3294-3301. |
| [15] | Zhuoran LI, Hua LI, Tong WANG, Chaozhe JIANG. Lightweight human pose estimation based on merge state space model [J]. Journal of Computer Applications, 2025, 45(10): 3179-3186. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||
