Journal of Computer Applications

    Next Articles

Music-driven visual semantic generation via dynamic music perception and cognitive reasoning

  

  • Received:2026-06-02 Revised:2026-08-15 Online:2026-09-09 Published:2026-09-09
  • About author:LI Dandan, born in 2004, M.S. candidate. Her research interests include multi‑modal large model. CHENG Ying, born in 1978, M.S., associate professor. Her research interests include higher music education. LEI Wenqiang, born in 1992, Ph.D., professor. His research interests include natural language processing.
  • Supported by:
    National Natural Science Foundation of China

基于动态音乐感知与认知推理的音乐驱动视觉语义生成方法

李丹丹1,承颖2,雷文强1   

  1. 1. 四川大学计算机学院
    2. 南京师范大学音乐学院
  • 通讯作者: 承颖
  • 作者简介:李丹丹(2004—),女,四川成都人,主要研究方向:多模态大模型;承颖(1978—),女,江苏南京人,副教授,硕士,主要研究方向:高等音乐教育;雷文强(1992—),男,四川威远人,教授,博士,主要研究方向:自然语言处理。
  • 基金资助:
    国家自然科学基金

Abstract: Existing music-driven visual semantic generation methods mostly rely on general music representations and open semantic reasoning, which suffer from insufficient alignment between musical information and visual semantic needs, as well as poor stability in cross-modal conversion. To address these issues, this paper proposes a music-driven visual semantic generation framework, MuViR. First, a dynamic music perception space is constructed, integrating four-dimensional perception attributesmusic energy, rhythmic motion, brightness, and complexityand their temporal variation characteristics to form a structured music representation oriented towards visual semantic expression needs, enhancing the dynamic expressive ability of music content in the cross-modal conversion process. Second, a Music Cognition Reasoning Chain (MCRC) is designed to decompose the music-to-visual conversion process hierarchically, guiding a large language model to sequentially complete music perceptual analysis, affective-state reasoning, and visual concept construction, establishing an intermediate conversion mechanism that aligns with music cognitive logic and improving the matching stability of audiovisual semantics. This paper conducts comparative validation on the Emotify dataset. Compared with From Sound to Sight (FSTS) and Generative Disco (GD), the proposed method achieves significant improvements in both CLAP score and semantic similarity. Experimental results show that acustomized music perceptual representation, combined with a hierarchical semantic inference mechanism, can effectively improve the consistency of music-visual semantic matching, providing reliable intermediate semantic support for downstream image and video generation tasks.

Key words: music representation, cross-modal conversion, music understanding, large language models, visual-semantic generation

摘要: 现有音乐驱动视觉语义生成方法多依赖通用音乐表征与开放式语义推理,存在音乐信息与视觉语义需求匹配度不足、跨模态转换稳定性差等问题。针对上述问题,本文提出一种音乐驱动视觉语义生成框架MuViR:首先,构建动态音乐感知空间,融合音乐能量、节奏运动、亮度、复杂性四维感知属性及其时序变化特征,形成面向视觉语义表达需求的结构化音乐表征,增强音乐内容在跨模态转换过程中的动态表达能力;其次,设计音乐认知推理链MCRC,对音乐到视觉的转换流程分层拆解,引导大语言模型依次完成音乐感知解析、情感状态推理和视觉概念构建,建立贴合音乐认知逻辑的中间转换机制,提升音视语义的匹配稳定性。本文基于Emotify数据集开展对比验证,相较于From Sound to Sight(FSTS)Generative Disco(GD),所提方法在CLAP分数、语义相似度两项指标上均取得明显提升。实验结果表明,定制化音乐感知表征结合分层语义推演机制,可有效提升音乐—视觉语义匹配一致性,为下游图像、视频生成任务提供可靠中间语义支撑。

关键词: 音乐表征, 跨模态转换, 音乐理解, 大语言模型, 视觉语义生成

CLC Number: