作者:Kejiang Xiao(肖克江), Cheng Zhang, Shiyan Pang*
出版刊物:Expert Systems with Applications
出版时间:2026年
内容摘要:
Multimodal Sentiment Analysis (MSA) aims to understand human emotions by integrating multiple modalities such as language, vision, and audio. However, existing methods typically rely on pairwise attention mechanisms, making it difficult to explicitly capture complex higher-order correlations and preserve fine-grained temporal cues during multimodal aggregation. To address these challenges, we propose the Hypergraph Dynamic Representation Network (HDR), a Structure-Sequence Co-modeling Framework for adaptive high-order multimodal interaction learning. Specifically, HDR constructs sample-specific dynamic hypergraphs through a coarse-to-fine topology learning strategy, where a K-Nearest Neighbor (K-NN) prior provides constrained semantic neighborhoods and a Gumbel-Softmax regulator adaptively refines the incidence structure. To mitigate temporal dilution, HDR further introduces a Dual Rotary Position Encoding (Dual-ROPE) mechanism that preserves temporal dependencies before and after hypergraph aggregation. Extensive experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS demonstrate that HDR consistently outperforms existing state-of-the-art methods, achieving an Acc-7 accuracy of 48.40% on CMU-MOSI. Furthermore, HDR exhibits clear efficiency advantages, reducing parameter overhead by approximately 72% while maintaining lower memory usage and fewer FLOPs. It also achieves low inference latency of 7.10ms/sample on CMU-MOSI and 1.69ms/sample on CMU-MOSEI, indicating that the dynamic hypergraph construction does not introduce prohibitive inference overhead.
多模态情感分析(MSA)旨在通过整合语言、视觉和音频等多种模态来理解人类情感。然而,现有方法通常依赖成对注意力机制,难以在多模态聚合过程中显式捕捉复杂的高阶关联并保留细粒度的时间线索。为解决上述挑战,我们提出超图动态表示网络(HDR),一种用于自适应高阶多模态交互学习的结构—序列协同建模框架。具体而言,HDR通过从粗到细的拓扑学习策略构建样本特定的动态超图,其中K近邻(K-NN)先验提供受约束的语义邻域,Gumbel-Softmax调节器自适应地细化关联结构。为缓解时间稀释,HDR进一步引入双旋转位置编码(Dual-ROPE)机制,在超图聚合前后保留时间依赖关系。在CMU-MOSI、CMU-MOSEI和CH-SIMS上的大量实验表明,HDR持续优于现有最先进方法,在CMU-MOSI上实现了48.40%的Acc-7准确率。此外,HDR展现出明显的效率优势,参数开销减少约72%,同时保持更低的内存占用和更少的FLOPs。它还在CMU-MOSI上实现了7.10ms/样本、在CMU-MOSEI上实现了1.69ms/样本的低推理延迟,表明动态超图构建不会引入过高的推理开销。