Journal of Computer Applications
Next Articles
Received:
Revised:
Accepted:
Online:
Published:
陈勇佳,李兴春*,程雨,何彦宇,李子亮
通讯作者:
Abstract: To address the issues of insufficient utilization of multi-level acoustic information, difficulties in cross-level feature alignment, and limited capability in characterizing temporal dynamics in Speech Emotion Recognition (SER), a speech emotion recognition method based on Dual-Channel Liquid Neural Network (DC-LNN) was proposed. Under a unified data augmentation strategy, two parallel branches were constructed. The Self-Supervised Separable Convolutional Neural Network (SSL-SCNN) branch was designed to extract high-level contextual features from self-supervised speech representations, and local temporal patterns were enhanced through depthwise separable convolutions. The Low-Level Descriptors-Liquid Neural Network (LLD-LNN) branch was constructed by taking Low-Level Descriptors (LLD), such as Mel-Frequency Cepstral Coefficients (MFCC), short-term energy, and zero-crossing rate, as inputs, and a Liquid Neural Network (LNN) was introduced to model the evolutionary process of emotions over time. On this basis, a Bi-Gated Cross-Hierarchical Attention (BG-CHA) fusion module was designed to achieve cross-branch alignment and adaptive information fusion. On the CASIA, EMODB, and SAVEE datasets, compared to the MSCRNN-A (Multi-Stream Convolutional Recurrent Neural Network based on Attention mechanism), AHPCL (Attention based Heterogeneous Parallel Convolutional Long short-term memory), and TIM-Net (Temporal-aware bI-direction Multi-scale Network), DC-LNN achieved average improvements of 18.83, 12.23, and 3.36 percentage points in Unweighted Accuracy Rate (UAR), respectively. This indicates that DC-LNN can effectively fuse acoustic information from different levels and improve speech emotion recognition performance. Ablation experiments and LLD branch temporal modeling comparison experiments further confirm that each core component contributes to the improvement of model performance. Experimental results show that DC-LNN can enhance the collaborative modeling ability of high- and low-level acoustic features and improve the stability of speech emotion recognition under complex acoustic conditions.
Key words: Speech Emotion Recognition (SER), Liquid Neural Network (LNN), dual-channel modeling, continuous-time modeling, feature fusion
摘要: 针对语音情感识别(SER)中多层声学信息利用不足、跨层级特征对齐困难和时序动态刻画能力有限的问题,提出一种基于双通道液态神经网络(DC-LNN)的语音情感识别方法。该方法在统一数据增强策略下构建两条并行分支:自监督深度可分离卷积神经网络(SSL-SCNN)分支基于自监督语音表征提取高层上下文特征,通过深度可分离卷积强化局部时序模式;低层描述符液态神经网络(LLD-LNN)分支以梅尔频率倒谱系数(MFCC)、短时能量和过零率等低层描述符(LLD)为输入,引入液态神经网络(LNN)对情绪随时间的演化过程进行建模。在此基础上,设计双向门控交叉层级注意力融合模块(BG-CHA),实现跨分支对齐与自适应信息融合。在CASIA、EMODB和SAVEE数据集上,相较于MSCRNN-A (Multi-Stream Convolutional Recurrent Neural Network based on Attention mechanism)、AHPCL (Attention based Heterogeneous Parallel Convolutional Long short-term memory)和TIM-Net (Temporal-aware bI-direction Multi-scale Network),DC-LNN在未加权准确率(UAR)指标上分别取得了18.83、12.23和3.36个百分点的平均提升,表明DC-LNN能够有效融合不同层级声学信息,提高语音情感识别性能;消融实验和LLD分支时序建模实验进一步证实,每个核心组件均有助于提升模型性能。实验结果表明,DC-LNN能够增强高层与低层声学特征的协同建模能力,提高复杂声学条件下语音情感识别的稳定性。
关键词: 语音情感识别, 液态神经网络, 双通道建模, 连续时间建模, 特征融合
CLC Number:
TP391.4
陈勇佳 李兴春 程雨 何彦宇 李子亮. 融合多层声学信息与液态神经网络的语音情感识别方法[J]. 《计算机应用》唯一官方网站, DOI: 10.11772/j.issn.1001-9081.2026040457.
/ Recommend
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2026040457