Journal of Computer Applications ›› 2026, Vol. 46 ›› Issue (9): 2761-2768.DOI: 10.11772/j.issn.1001-9081.2025080959

• Artificial intelligence • Previous Articles    

Multimodal sentiment analysis model for missing modalities under shared semantic conditions

Shang LIU(), Zhaosen TANG, Hongyue LIU, Linfang DONG, Jin ZHOU   

  1. School of Science and Technology,Tianjin University of Finance and Economics,Tianjin 300222,China
  • Received:2025-08-22 Revised:2025-10-03 Accepted:2025-10-14 Online:2026-09-16 Published:2026-09-10
  • Contact: Shang LIU
  • About author:LIU Shang,born in 1977, Ph. D., professor. Her research interestsinclude pattern recognition, digital image processing.
    TANG Zhaosen, born in 2002, M. S. candidate. His researchinterests include pattern recognition, multimodal perception.
    LIU Hongyue, born in 2001, M. S. candidate. Her researchinterests include image processing, pose estimation.
    DONG Linfang, born in 1972, Ph. D., associate professor. Herresearch interests include artificial intelligence, natural language processing.
    ZHOU Jin, born in 1981, Ph. D., associate professor. Her researchinterests include modulation recognition and spectrum sensing based ondeep learning.

共享语义条件下面向模态缺失的多模态情感分析模型

刘赏(), 汤兆森, 刘鸿月, 董林芳, 周金   

  1. 天津财经大学 理工学院,天津 300222
  • 通讯作者: 刘赏
  • 作者简介:刘赏(1977—),女,河北辛集人,教授,博士,主要研究方向:模式识别、数字图像处理
    汤兆森(2002—),男,河南濮阳人,硕士研究生,主要研究方向:模式识别、多模态感知
    刘鸿月(2001—),女,贵州贵阳人,硕士研究生,主要研究方向:图像处理、姿态估计
    董林芳(1972—),女,河北张家口人,副教授,博士,主要研究方向:人工智能、自然语言处理
    周金(1981—),女,天津人,副教授,博士,主要研究方向:基于深度学习的调制识别及频谱感知。
  • 基金资助:
    天津市自然科学基金资助项目(22JCYBJC01550);天津市艺术科学规划项目(C22030)

Abstract:

Existing studies on multimodal sentiment analysis in modality missing scenarios often neglect inter-modal correlations when generating missing modalities, leading to semantic inconsistencies between restored and original data. Additionally, missing modality generation based on diffusion models brings large computational overhead. To address these issues, a Multimodal Sentiment Analysis Model for Missing Modalities under Shared Semantic Conditions (MM-SSC) was proposed. First, a shared latent space mapping module was designed to use the shared latent space of Vector Quantized Variational AutoEncoder (VQ-VAE) to capture multimodal distributions effectively, thereby ensuring cross-modal shared semantics. Second, a cross-modal consistency constraint method was proposed to learn mutual information within each modality's latent space, thereby promoting the refinement of semantic information between modalities and enhancing cross-modal consistency. Third, a missing modality reconstruction and alignment module was designed to reconstruct and refine missing modalities while reducing reconstruction computational overhead. Finally, a multimodal fusion and prediction module was introduced to fuse reconstructed and available modalities for consistent sentiment analysis. Experimental results demonstrate that under fixed modality missing conditions, compared with Incomplete Multimodality-Diffused emotion recognition (IMDer) model, the proposed model achieves average improvements of 0.4 and 0.7 percentage points in the F1 score and ACC7 on the CMU-MOSI dataset, respectively; on the CMU-MOSEI dataset, the F1 score and ACC7 are improved by 0.9 and 0.3 percentage points, respectively. Under random modality missing conditions, compared with IMDer model, the proposed model achieves an average improvement of 2.2 percentage points in both F1 score and ACC7 on the CMU-MOSI dataset; on the CMU-MOSEI dataset, the F1 score and ACC7 (accuracy for seven classes) are improved by 0.8 and 0.4 percentage points, respectively. It can be seen that MM-SSC can address multimodal sentiment analysis tasks under modality missing scenarios effectively.

Key words: multimodal sentiment analysis, missing modality, deep learning, Vector Quantized Variational AutoEncoder (VQ-VAE), diffusion model

摘要:

现有研究在进行模态缺失场景下的多模态情感分析时,在缺失模态生成过程中常忽略模态间的相关性,导致恢复数据与原始数据的语义不一致,同时基于扩散模型的缺失模态生成会带来较大的计算开销。针对以上问题,提出一种共享语义条件下面向模态缺失的多模态情感分析模型(MM-SSC)。首先,设计共享潜空间映射模块,利用矢量量化变分自编码器(VQ-VAE)的共享潜空间有效捕捉多模态分布,确保模态间共享语义;其次,给出跨模态一致性约束方法,以学习每种模态潜空间内的互信息,促进模态间语义信息的完善,增强跨模态一致性;再次,设计缺失模态重建与对齐模块,以重建与细化缺失模态,同时降低重建计算开销;最后,引入多模态融合与预测模块,融合重建模态与可用模态,用于一致性的情感分析。实验结果显示,在固定模态缺失的情况下,相较于不完全多模态扩散情感识别(IMDer)模型,所提模型在CMU-MOSI数据集上的F1得分和七分类准确率(ACC7)指标分别平均提升了0.4与0.7个百分点,在CMU-MOSEI数据集上的F1得分和ACC7分别平均提升了0.9与0.3个百分点;在随机模态缺失的情况下,相比IMDer模型,所提模型在CMU-MOSI数据集上的F1得分和ACC7都平均提升了2.2个百分点,在CMU-MOSEI数据集上的F1得分和ACC7分别平均提升了0.8与0.4个百分点。MM-SSC能有效应对模态缺失场景下的多模态情感分析任务。

关键词: 多模态情感分析, 模态缺失, 深度学习, 矢量量化变分自编码器, 扩散模型

CLC Number: