《计算机应用》唯一官方网站 ›› 2026, Vol. 46 ›› Issue (8): 2515-2523.DOI: 10.11772/j.issn.1001-9081.2025070826
吕仁堃1, 孙鹏1,2(
), 郎宇博1, 郭弘2, 沈喆3, 田迪4
收稿日期:2025-07-23
修回日期:2025-09-25
接受日期:2025-09-26
发布日期:2025-11-05
出版日期:2026-08-10
通讯作者:
孙鹏
作者简介:吕仁堃(2000—),男,山东烟台人,硕士研究生,CCF会员,主要研究方向:深度伪造检验、图像处理基金资助:
Renkun LYU1, Peng SUN1,2(
), Yubo LANG1, Hong GUO2, Zhe SHEN3, Di TIAN4
Received:2025-07-23
Revised:2025-09-25
Accepted:2025-09-26
Online:2025-11-05
Published:2026-08-10
Contact:
Peng SUN
About author:LYU Renkun, born in 2000, M. S. candidate. His research interests include deepfake detection, image processing.Supported by:摘要:
现有深度伪造检验方法多数依据图像像素级线索建模,较少考虑合成过程对伪造图像的影响,尽管取得较好的检验效果,也很难解释检验过程。因此,提出一种多模态物理先验特征融合的深度伪造可解释检验方法。首先,使用光流特征、光照特征、边缘特征和离散余弦变换(DCT)特征分别描述时序视频中的帧间运动差异性、单帧视频中的光照不一致性和边缘伪影信息,得到具有可解释性的多模态物理先验特征;其次,提出多模态混合专家网络,分别构建不同模态的专家子网络,并把子网络在跨模态注意力加权后经门控单元融合,并输入判别网络中实现分类;然后,在判别网络中引入SIAM(Spatial Intersection Attention Module),并将全连接结构替换为KAN(Kolmogorov-Arnold Network)结构;最后,利用多模态物理先验特征分别对不同的专家子网络进行训练,给出对不同输入特征的Shapley值分析,构建事前特征-事后解释的可解释分析框架,为模型推理和预测提供像素级解释。实验结果表明,与CORE(COnsistent REpresentation learning)、SRM(Rich Models for Steganalysis)和UCF(Uncovering Common Features)等算法相比,所提方法在FaceForensics++数据集上的ROC曲线下面积(AUC)和准确率最好,准确率范围为97.35%~98.75%,平均准确率达98.22%,模型可解释性也有较大提升。
中图分类号:
吕仁堃, 孙鹏, 郎宇博, 郭弘, 沈喆, 田迪. 多模态物理先验特征融合的深度伪造检验方法[J]. 计算机应用, 2026, 46(8): 2515-2523.
Renkun LYU, Peng SUN, Yubo LANG, Hong GUO, Zhe SHEN, Di TIAN. Deepfake detection method based on fusion of multi-modal physical prior features[J]. Journal of Computer Applications, 2026, 46(8): 2515-2523.
| 方法 | DF | F2F | FS | NT | ||||
|---|---|---|---|---|---|---|---|---|
| ACC/% | AUC | ACC/% | AUC | ACC/% | AUC | ACC/% | AUC | |
| Xception | 96.58 | 0.979 9 | 96.70 | 0.978 5 | 97.23 | 0.983 3 | 92.72 | 0.938 5 |
| Capsule | 86.02 | 0.866 9 | 85.41 | 0.863 4 | 86.29 | 0.873 4 | 75.95 | 0.780 4 |
| Face X-ray | 96.51 | 0.979 4 | 97.73 | 0.987 2 | 97.68 | 0.987 1 | 91.84 | 0.929 0 |
| CORE | 96.49 | 0.978 7 | 97.22 | 0.980 3 | 97.35 | 0.982 3 | 92.30 | 0.933 9 |
| UCF | 98.15 | 0.988 3 | 97.53 | 0.984 0 | 98.07 | 0.989 6 | 93.37 | 0.944 1 |
| F3Net | 96.51 | 0.979 3 | 96.86 | 0.979 6 | 97.58 | 0.984 4 | 95.44 | 0.965 4 |
| SRM | 96.87 | 0.973 3 | 95.95 | 0.969 6 | 96.41 | 0.974 4 | 91.39 | 0.924 5 |
| 本文方法 | 98.75 | 0.998 7 | 98.11 | 0.992 1 | 98.66 | 0.996 9 | 97.35 | 0.983 5 |
表1 不同方法检测结果对比
Tab. 1 Comparison of detection results using different methods
| 方法 | DF | F2F | FS | NT | ||||
|---|---|---|---|---|---|---|---|---|
| ACC/% | AUC | ACC/% | AUC | ACC/% | AUC | ACC/% | AUC | |
| Xception | 96.58 | 0.979 9 | 96.70 | 0.978 5 | 97.23 | 0.983 3 | 92.72 | 0.938 5 |
| Capsule | 86.02 | 0.866 9 | 85.41 | 0.863 4 | 86.29 | 0.873 4 | 75.95 | 0.780 4 |
| Face X-ray | 96.51 | 0.979 4 | 97.73 | 0.987 2 | 97.68 | 0.987 1 | 91.84 | 0.929 0 |
| CORE | 96.49 | 0.978 7 | 97.22 | 0.980 3 | 97.35 | 0.982 3 | 92.30 | 0.933 9 |
| UCF | 98.15 | 0.988 3 | 97.53 | 0.984 0 | 98.07 | 0.989 6 | 93.37 | 0.944 1 |
| F3Net | 96.51 | 0.979 3 | 96.86 | 0.979 6 | 97.58 | 0.984 4 | 95.44 | 0.965 4 |
| SRM | 96.87 | 0.973 3 | 95.95 | 0.969 6 | 96.41 | 0.974 4 | 91.39 | 0.924 5 |
| 本文方法 | 98.75 | 0.998 7 | 98.11 | 0.992 1 | 98.66 | 0.996 9 | 97.35 | 0.983 5 |
| 方法 | 参数量/106 | GFLOPs | 方法 | 参数量/106 | GFLOPs |
|---|---|---|---|---|---|
| Xception | 21.86 | 16.71 | UCF | 46.95 | 24.60 |
| Capsule | 38.98 | 32.43 | F3Net | 22.57 | 12.84 |
| Face X-ray | 77.57 | 19.50 | SRM | 55.47 | 25.23 |
| CORE | 21.92 | 12.25 | 本文方法 | 23.35 | 8.86 |
表2 不同方法的参数量和浮点运算量对比
Tab. 2 Comparison of parameters and floating point operations for different methods
| 方法 | 参数量/106 | GFLOPs | 方法 | 参数量/106 | GFLOPs |
|---|---|---|---|---|---|
| Xception | 21.86 | 16.71 | UCF | 46.95 | 24.60 |
| Capsule | 38.98 | 32.43 | F3Net | 22.57 | 12.84 |
| Face X-ray | 77.57 | 19.50 | SRM | 55.47 | 25.23 |
| CORE | 21.92 | 12.25 | 本文方法 | 23.35 | 8.86 |
| 模型 | ACC/% | AUC |
|---|---|---|
| 92.65 | 0.980 1 | |
| 97.42 | 0.993 5 | |
| 97.96 | 0.996 0 | |
| 本文模型 | 98.75 | 0.998 7 |
表3 不同模态融合消融实验结果对比
Tab. 3 Comparison of ablation experimental results with different modal fusions
| 模型 | ACC/% | AUC |
|---|---|---|
| 92.65 | 0.980 1 | |
| 97.42 | 0.993 5 | |
| 97.96 | 0.996 0 | |
| 本文模型 | 98.75 | 0.998 7 |
| Backbone | SIAM | KAN | ACC/% | AUC |
|---|---|---|---|---|
| √ | 98.08 | 0.995 8 | ||
| √ | √ | 98.12 | 0.997 4 | |
| √ | √ | √ | 98.75 | 0.998 7 |
表4 不同改进方案消融实验结果对比
Tab. 4 Comparison of ablation experimental results with different improvements
| Backbone | SIAM | KAN | ACC/% | AUC |
|---|---|---|---|---|
| √ | 98.08 | 0.995 8 | ||
| √ | √ | 98.12 | 0.997 4 | |
| √ | √ | √ | 98.75 | 0.998 7 |
| [1] | Amerini I, Galteri L, Caldelli R, et al. Deepfake video detection through optical flow based CNN[C]// ICCVW 2019. Piscataway: IEEE, 2019: 1205-1207. |
| [2] | Sun D, Yang X, Liu M Y, et al. PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume[C]// CVPR 2018. Piscataway: IEEE, 2018: 8934-8943. |
| [3] | 吴文轩,周文柏,张卫明,等. 基于块间光照不一致性的深度伪造检测算法[J]. 网络与信息安全学报, 2023, 9(1): 167-177. |
| Wu Wenxuan, Zhou Wenbo, Zhang Weiming, et al. Deepfake detection method based on patch-wise lighting inconsistency[J]. Chinese Journal of Network and Information Security, 2023, 9(1): 167-177. | |
| [4] | Liu Z, Wang Y, Vaidya S, et al. KAN: Kolmogorov-Arnold networks[PP/OL]. V5. arXiv (2025-02-09) [2025-06-20].. |
| [5] | 李纪成,刘琲贝,胡永健,等. 基于光照方向一致性的换脸视频检测[J]. 南京航空航天大学学报, 2020, 52(5): 760-767. |
| Li Jicheng, Liu Beibei, Hu Yongjian, et al. Deepfake video detection based on consistency of illumination direction[J]. Journal of Nanjing University of Aeronautics and Astronautics, 2020, 52(5): 760-767. | |
| [6] | 杨珂,李永亮,何金栋,等. 基于掩码图像建模的深度伪造人脸检测[J]. 计算机应用, 2025, 45(S1): 72-77. |
| Yang Ke, Li Yongliang, He Jindong, et al. Deepfake face detection based on masked image modeling[J]. Journal of Computer Applications, 2025, 45(S1): 72-77. | |
| [7] | Fu X, Fu B, Chen S, et al. Faces blind your eyes: unveiling the content-irrelevant synthetic artifacts for deepfake detection[J]. IEEE Transactions on Image Processing, 2025, 34: 5686-5696. |
| [8] | Li L, Bao J, Zhang T, et al. Face X-ray for more general face forgery detection[C]// CVPR 2020. Piscataway: IEEE, 2020: 5000-5009. |
| [9] | Dong X, Bao J, Chen D, et al. Protecting celebrities from deepfake with identity consistency Transformer[C]// CVPR 2022. Piscataway: IEEE, 2022: 9458-9468. |
| [10] | Ganguly S, Ganguly A, Mohiuddin S, et al. ViXNet: vision Transformer with Xception network for deepfakes based video and image forgery detection[J]. Expert Systems with Applications, 2022, 210: No.118423. |
| [11] | Qian Y, Yin G, Sheng L, et al. Thinking in frequency: face forgery detection by mining frequency-aware clues[C]// ECCV 2020, LNCS 12357. Cham: Springer, 2020: 86-103. |
| [12] | Cheng Z, Wang Y, Wan Y, et al. DeepFake detection method based on multi-scale interactive dual-stream network[J]. Journal of Visual Communication and Image Representation, 2024, 104: No.104263. |
| [13] | Wang J, Wu Z, Ouyang W, et al. M2TR: multi-modal multi-scale transformers for deepfake detection[C]// ICMR 2022. New York: ACM, 2022: 615-623. |
| [14] | Cozzolino D, Rössler A, Thies J, et al. ID-Reveal: identity-aware deepfake video detection[C]// ICCV 2021. Piscataway: IEEE, 2021: 15088-15097. |
| [15] | Conotter V, Bodnari E, Boato G, et al. Physiologically-based detection of computer generated faces in video[C]// ICIP 2014. Piscataway: IEEE, 2014: 248-252. |
| [16] | Ciftci U A, Demir I, Yin L. FakeCatcher: detection of synthetic portrait videos using biological signals[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020(Early Access): 1. |
| [17] | Li Y, Chang M C, Lyu S. In ictu oculi: exposing AI created fake videos by detecting eye blinking[C]// WIFS 2018. Piscataway: IEEE, 2018: 1-7. |
| [18] | Sun Z, Han Y, Hua Z, et al. Improving the efficiency and robustness of deepfakes detection through precise geometric features[C]// CVPR 2021. Piscataway: IEEE, 2021: 3608-3617. |
| [19] | Cozzolino D, Pianese A, Nießner M, et al. Audio-visual person-of-interest deepfake detection[C]// CVPR 2023. Piscataway: IEEE, 2023: 943-952. |
| [20] | 肖景博,殷琪林,卢伟,等. 基于视频流谱特征空间的深度伪造检测[J]. 中国科学:信息科学, 2024, 54(11): 2572-2588. |
| Xiao Jingbo, Yin Qilin, Lu Wei, et al. Deepfake detection based on video flow spectrum feature space[J]. SCIENTIA SINICA Informationis, 2024, 54(11): 2572-2588. | |
| [21] | Zhou Y, Lim S N. Joint audio-visual deepfake detection[C]// ICCV 2021. Piscataway: IEEE, 2021: 1480-1489. |
| [22] | Cai Z, Ghosh S, Dhall A, et al. Glitch in the matrix: a large scale benchmark for content driven audio-visual forgery detection and localization[J]. Computer Vision and Image Understanding, 2023, 236: No.103818. |
| [23] | Jacobs R A, Jordan M I, Nowlan S J, et al. Adaptive mixtures of local experts[J]. Neural Computation, 1991, 3(1): 79-87. |
| [24] | Negroni V, Salvi D, Mezza A I, et al. Leveraging mixture of experts for improved speech deepfake detection[C]// ICASSP 2025. Piscataway: IEEE, 2025: 1-5. |
| [25] | Shazeer N, Mirhoseini A, Maziarz K, et al. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer[EB/OL]. (2025-10-12) [2025-10-20].. |
| [26] | Rozemberczki B, Watson L, Bayer P, et al. The shapley value in machine learning[C]// IJCAI 2022. California: IJCAI, 2022: 5572-5579. |
| [27] | Han G, Huang S, Zhao F, et al. SIAM: a parameter-free, spatial intersection attention module[J]. Pattern Recognition, 2024, 153: No.110509. |
| [28] | Sun D, Roth S, Black M J. Secrets of optical flow estimation and their principles[C]// CVPR 2010. Piscataway: IEEE, 2010: 2432-2439. |
| [29] | Huang Z, Shi X, Zhang C, et al. FlowFormer: a Transformer architecture for optical flow[C]// ECCV 2022, LNCS 13677. Cham: Springer, 2022: 668-685. |
| [30] | Yao L, Lin Y, Muhammad S. An improved multi-scale image enhancement method based on retinex theory[J]. Journal of Medical Imaging and Health Informatics, 2018, 8(1): 122-126. |
| [31] | Zhang H, Hu C, Min S, et al. TSFF-Net: a deep fake video detection model based on two-stream feature domain fusion[J]. PLoS ONE, 2024, 19(12): No.e0311366. |
| [32] | Rössler A, Cozzolino D, Verdoliva L, et al. FaceForensics++: learning to detect manipulated facial images[C]// ICCV 2019. Piscataway: IEEE, 2019: 1-11. |
| [33] | Chollet F. Xception: deep learning with depthwise separable convolutions[C]// CVPR 2017.Piscataway: IEEE,2017:1800-1807. |
| [34] | Nguyen H H, Yamagishi J, Echizen I. Capsule-forensics: using capsule networks to detect forged images and videos[C]// ICASSP 2019. Piscataway: IEEE, 2019: 2307-2311. |
| [35] | Ni Y, Meng D, Yu C, et al. CORE: consistent representation learning for face forgery detection[C]// CVPR 2022. Piscataway: IEEE, 2022: 12-21. |
| [36] | Yan Z, Zhang Y, Fan Y, et al. UCF: uncovering common features for generalizable deepfake detection[C]// ICCV 2023. Piscataway: IEEE, 2023: 22355-22366. |
| [37] | Luo Y, Zhang Y, Yan J, et al. Generalizing face forgery detection with high-frequency features[C]// CVPR 2021. Piscataway: IEEE, 2021: 16312-16321. |
| [1] | 邱原, 彭海龙, 费蓉, 徐庆征, 李仟禧, 薛诚. 多视图注意力融合的引文网络学术群体识别[J]. 《计算机应用》唯一官方网站, 2026, 46(8): 2494-2504. |
| [2] | 惠康华 于慕涵 张智 赵敏. 基于跨域泛化增强的立体匹配方法[J]. 《计算机应用》唯一官方网站, 0, (): 0-0. |
| [3] | 李建东 王圣龙 折佳佳 齐潺彬. 融合特征增强和注意力优化的复杂环境水下目标检测[J]. 《计算机应用》唯一官方网站, 0, (): 0-0. |
| [4] | 马岳, 赖惠成, 姜迪, 汪烈军. 多模态协同提示优化下的零样本人物交互检测方法[J]. 《计算机应用》唯一官方网站, 2026, 46(7): 2277-2287. |
| [5] | 杜秀丽, 高星, 张校毓, 潘成胜, 邹启杰. 基于密集时空可变形注意力的视频快照压缩成像重建方法[J]. 《计算机应用》唯一官方网站, 2026, 46(7): 2288-2296. |
| [6] | 张国有, 聂宏宇, 潘理虎, 雷润东. 基于多层感知机级联宽度学习系统的点云语义分割网络Point-MLPBLS[J]. 《计算机应用》唯一官方网站, 2026, 46(7): 2259-2266. |
| [7] | 颜建强, 董贝贝, 曲博婷, 彭晨. 融合多源信息与图级注意力的双向扩散动态图卷积交通流预测网络[J]. 《计算机应用》唯一官方网站, 2026, 46(7): 2327-2333. |
| [8] | 刘苗苗, 张郁红, 张强, 杜睿山. 基于改进GhostNet的岩石薄片岩性识别模型[J]. 《计算机应用》唯一官方网站, 2026, 46(7): 2355-2363. |
| [9] | 黄宝来 曾维 朱星 潘玉杰. 面向高分辨率图像去模糊的跨域多尺度网络[J]. 《计算机应用》唯一官方网站, 0, (): 0-0. |
| [10] | 曹型兵 张翔 李校林. 面向路面病害检测的边缘频率感知网络EFA-Net[J]. 《计算机应用》唯一官方网站, 0, (): 0-0. |
| [11] | 武俊丽 李文欣 李建辉. 基于自适应特征过滤与重组的小样本表面缺陷分类[J]. 《计算机应用》唯一官方网站, 0, (): 0-0. |
| [12] | 吕超, 马歌谣. 基于冗余特征抑制的轻量级人体姿态估计网络[J]. 《计算机应用》唯一官方网站, 2026, 46(6): 1973-1980. |
| [13] | 舒尔豪, 涂国庆, 刘树波. 基于类激活映射的空频协同对抗样本生成方法[J]. 《计算机应用》唯一官方网站, 2026, 46(6): 1881-1892. |
| [14] | 张纾豪, 何坤金, 徐佳晨, 沙河山, 陈正鸣. 融合透视校正与轻量注意力机制的轮毂缺陷检测方法[J]. 《计算机应用》唯一官方网站, 2026, 46(6): 2007-2015. |
| [15] | 张金萧, 李成龙, 高新燕, 张铭. 基于时空特征金字塔网络与多假设交互机制的三维人体姿态估计模型[J]. 《计算机应用》唯一官方网站, 2026, 46(6): 1965-1972. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||