Journals
  Publication Years
  Keywords
Search within results Open Search
Please wait a minute...
For Selected: Toggle Thumbnails
Deepfake detection method based on fusion of multi-modal physical prior features
Renkun LYU, Peng SUN, Yubo LANG, Hong GUO, Zhe SHEN, Di TIAN
Journal of Computer Applications    2026, 46 (8): 2515-2523.   DOI: 10.11772/j.issn.1001-9081.2025070826
Abstract155)   HTML1)    PDF (2963KB)(18)       Save

The existing deepfake detection methods mainly model on the basis of pixel-level clues of images, and seldom consider the impact of the synthesis process on the forged images. Although good detection results are achieved, it is difficult to explain the detection process. Therefore, a multi-modal physical prior feature fusion-based explainable detection method for deepfakes was proposed. First, optical flow features, illumination features, edge features and DCT(Discrete Cosine Transform) features were used to describe the inter-frame motion differences in temporal videos, the illumination inconsistency in single-frame videos, and edge artifact information, respectively, so as to obtain multi-modal physical prior features with explainability. Second, a multi-modal mixture of experts network was proposed to construct expert sub-networks for different modalities, and after cross-modal attention weighting, the sub-networks were fused through a gated unit and input into the discriminative network for classification. Third, the SIAM (Spatial Intersection Attention Module) was introduced into the discriminative network, and the fully connected structure was replaced by the KAN (Kolmogorov-Arnold Network) structure. Finally, the multi-modal physical prior features were used to train different expert sub-networks, respectively, and the Shapley value analysis of different input features was given, thereby constructing a pre-feature-post-explanation explainable analysis framework to provide pixel-level explanations for model inference and prediction. Experimental results show that compared with algorithms such as CORE(COnsistent REpresentation learning), SRM(Rich Models for Steganalysis), and UCF(Uncovering Common Features), the proposed method achieves the best performance on AUC and accuracy, with an accuracy range of 97.35% to 98.75% and an average accuracy of 98.22% on FaceForensics++ dataset, and the model’s interpretability also is improved significantly.

Table and Figures | Reference | Related Articles | Metrics