Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (7): 2111-2118.DOI: 10.11772/j.issn.1001-9081.2025070906

• Artificial intelligence • Previous Articles    

Interpretation method based on selective Softmax gradient for layer-wise relevance propagation

Chong CHEN1, Hongyang ZOU1, Jing CAO1, Wenjing JIANG1, Jie CHEN1, Fumin GAO2()   

  1. 1.College of Artificial Intelligence,China University of Petroleum (Beijing),Beijing 102249,China
    2.College of Safety and Ocean Engineering,China University of Petroleum (Beijing),Beijing 102249,China
  • Received:2025-08-11 Revised:2025-11-06 Accepted:2025-11-10 Online:2025-12-22 Published:2026-07-10
  • Contact: Fumin GAO
  • About author:CHEN Chong, born in 1987, Ph. D., associate professor. His research interests include machine learning and its interpretability, numerical modeling.
    ZOU Hongyang, born in 2001, M. S. candidate. His research interests include interpretability of neural network models.
    CAO Jing, born in 2002, M. S. candidate. His research interests include interpretability of neural network models.
    JIANG Wenjing, born in 2002, M. S. Her research interests include data assimilation.
    CHEN Jie, born in 1997, M. S. Her research interests include interpretability of deep learning models.
  • Supported by:
    National Key Research and Development Program of China(2022YFC2803700);Oil&Gas Major Project(2024ZD1403305)

基于选择性Softmax梯度的分层相关性传播解释方法

陈冲1, 邹宏杨1, 曹靖1, 蒋文静1, 陈杰1, 高富民2()   

  1. 1.中国石油大学(北京) 人工智能学院,北京 102249
    2.中国石油大学(北京) 安全与海洋工程学院,北京 102249
  • 通讯作者: 高富民
  • 作者简介:陈冲(1987—),男,河北衡水人,副教授,博士,CCF会员,主要研究方向:机器学习及其可解释性、数值模拟
    邹宏杨(2001—),男,四川成都人,硕士研究生,主要研究方向:神经网络模型的可解释性
    曹靖(2002—),男,北京人,硕士研究生,主要研究方向:神经网络模型的可解释性
    蒋文静(2002—),女,湖南常德人,硕士,主要研究方向:数据同化
    陈杰(1997—),女,山东菏泽人,硕士,主要研究方向:深度学习可解释性
  • 基金资助:
    国家重点研发计划项目(2022YFC2803700);油气重大专项(2024ZD1403305)

Abstract:

The inherent “black-box” nature of neural network models leads to opaque internal logic and behavioral decision, restricting their application in critical areas. To address this issue, based on the Layer-wise Relevance Propagation (LRP) method, the interpretability of image classification models such as VGG16 and ResNet50 was researched, and an interpretation method for image classification models named SSGLRP (Selective Softmax Gradient for Layer-wise Relevance Propagation) was proposed. By introducing activation values for the positive gradients of output neurons and modifying the initial relevance values for non-target classes, the problem of LRP heatmaps containing noise and lacking class discrimination was effectively solved. The SSGLRP method was quantitatively evaluated using maximum patch masking and fixed-point game experiments, with classic LRP, Selective Layer-wise Relevance Propagation (SLRP), and Softmax Gradient Layer-wise Relevance Propagation (SGLRP) as baseline methods. The results of the maximum patch masking experiments show that on VGG16, the prediction change obtained by SSGLRP is on average 91.8%, 60.3%, and 34.2% higher than those obtained by LRP, SLRP, and SGLRP, respectively; on ResNet50, the improvements are 61.9%, 64.9%, and 33.2%, respectively. The SSGLRP method has higher class discrimination ability and less noise with superior performance in interpreting VGG16 and ResNet50 models.

Key words: neural network, interpretation method, class discrimination, Layer-wise Relevance Propagation (LRP), image classification

摘要:

神经网络模型固有的“黑盒”性质导致它们的内部逻辑与行为决策不透明,限制了它们在关键领域的应用。针对该问题,基于分层相关传播(LRP)方法,对VGG16和ResNet50等图像分类模型的可解释性展开研究,并提出一种针对图像分类模型的可解释性方法SSGLRP (Selective Softmax Gradient for Layer-wise Relevance Propagation)。该方法在LRP中引入输出神经元正梯度的激活值与修正非目标类的初始相关性值,解决了LRP方法解释结果的热力图中含有噪声且缺乏类别区分性问题。以经典LRP、选择性分层相关传播(SLRP)、Softmax梯度分层相关传播(SGLRP)为基线方法,采用最大补丁遮掩和定点游戏实验对SSGLRP方法进行定量评估。最大补丁遮掩的实验结果表明,在VGG16上,SSGLRP的预测值变化量比LRP、SLRP和SGLRP分别平均高出91.8%、60.3%和34.2%;在ResNet50上则分别高出61.9%、64.9%和33.2%。SSGLRP方法具有更高的类别区别能力和更少的噪声,且解释VGG16和ResNet50模型的效果更好。

关键词: 神经网络, 解释方法, 类别区分性, 分层相关传播, 图像分类

CLC Number: