A Dual-Stream Temporal Enhancement and Contribution Weight Fusion Model for Sentiment Analysis
| 62 | 0 | 147 |
| 下载次数 | 被引频次 | 阅读次数 |
针对现有多模态情感分析方法难以捕捉长文本语义、视频时序动态演变不足以及图文融合方式不合理的问题,提出一种用于人物对话视频情感分析场景的双流协同处理模型DREA(dual stream-RLTBERTa-E-HTViT-attention)。该模型包括RLTBERTa(RoBERTa with LSTM and Transformer bidirectional encoder representations from Transformers with attention)文本流与E-HTViT(emotion-aware hierarchical temporal vision Transformer)视觉流两个并行模块,文本流通过融合双向长短期记忆网络(BiLSTM)与Transformer结构增强长距离语义建模能力,视觉流通过层次化时序编码捕捉短时微表情与长时情感演化;采用贡献权重融合机制实现双模态表征的动态自适应融合。实验结果表明,与现有主流多模态模型相比较,本文模型在人物对话视频情感分析任务中表现出较好的识别性能,在影视内容分析与人机情感交互等场景中具有一定的应用潜力。
Abstract:To address the limitations of existing multimodal sentiment analysis methods in capturing long-text semantics, insufficiently modeling the temporal dynamic evolution of videos, and adopting unreasonable image-text fusion strategies, a dual-stream collaborative processing model named DREA (dual stream-RLTBERTa-E-HTViT-attention) is proposed for character dialogue video sentiment analysis. The model consists of two parallel modules: the RLTBERTa (RoBERTa with LSTM and Transformer bidirectional encoder representations from Transformers with attention) text stream and the E-HTViT (emotion-aware hierarchical temporal vision Transformer) visual stream. In the text stream, the bidirectional long short-term memory network (BiLSTM) and Transformer structure are integrated to enhance long-range semantic modeling capability. In the visual stream, hierarchical temporal encoding is employed to capture short-term micro-expressions and long-term emotional evolution. In addition, a contribution-weighted fusion mechanism is adopted to achieve dynamic and adaptive fusion of bimodal representations. Experimental results show that, compared with existing mainstream multimodal models, the proposed model demonstrates favorable recognition performance in character dialogue video sentiment analysis tasks and exhibits application potential in film and television content analysis, human-computer emotional interaction, and other related scenarios.
[1] Zhao L G, Lee S W. Integrating ontology-based approaches with deep learning models for fine-grained sentiment analysis[J]. Computers, Materials & Continua, 2024, 81(1): 1855-1877.
[2] Das R, Singh T D. Multimodal sentiment analysis: a survey of methods, trends, and challenges[J]. ACM Computing Surveys, 2023, 55(13s): 1-38.
[3] Ortis A, Farinella G M, Torrisi G, et al. Exploiting objective text description of images for visual sentiment analysis[J]. Multimedia Tools and Applications, 2021, 80(15): 22323-22346.
[4] Yang J F, She D Y, Lai Y K, et al. Retrieving and classifying affective images via deep metric learning[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2018, 32(1): 491-498.
[5] Sai Raviteja Chappa N V, Luu K. LiGAR: LiDAR-guided hierarchical transformer for multi-modal group activity recognition[C]//2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson, AZ, USA: IEEE, 2025: 3035-3044.
[6] Yang L. A dynamic weighted fusion model for multimodal sentiment analysis[J]. Signal, Image and Video Processing, 2025, 19(8): 609.
[7] Cui Y M, Che W X, Liu T, et al. Pre-training with whole word masking for Chinese BERT[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3504-3514.
[8] Wang N, Wang Q. Dynamic weighted gating for enhanced cross-modal interaction in multimodal sentiment analysis[J]. ACM Transactions on Multimedia Computing, Communications, and Applications, 2025, 21(1): 1-19.
[9] Afroze S, Hossain M R, Hoque M M, et al. MulMoSenT: multimodal sentiment analysis for a low-resource language using textual-visual cross-attention and fusion[J]. Information Fusion, 2026, 131: 104129.
[10] Miao Y L, Cheng W F, Ji Y C, et al. Aspect-based sentiment analysis in Chinese based on mobile reviews for BiLSTM-CRF[J]. Journal of Intelligent & Fuzzy Systems, 2021, 40(5): 8697-8707.
[11] Sornlertlamvanich V, Yuenyong S. Thai named entity recognition using BiLSTM-CNN-CRF enhanced by TCC[J]. IEEE Access, 2022, 10: 53043-53052.
[12] Devlin J, Chang M W, Lee K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, Minnesota, USA: ACL, 2019: 4171-4186.
[13] Liu N, Hu Q, Xu H Y, et al. Med-BERT: a pretraining framework for medical records named entity recognition[J]. IEEE Transactions on Industrial Informatics, 2022, 18(8): 5600-5608.
[14] Zhou H, Tang S L, Huang W, et al. Generating risk response measures for subway construction by fusion of knowledge and deep learning[J]. Automation in Construction, 2023, 152: 104951.
[15] Xu Y J, Tan X B, Tong X, et al. A robust Chinese named entity recognition method based on integrating dual-layer features and CSBERT[J]. Applied Sciences, 2024, 14(3): 1060.
[16] Urolagin S, Nayak J, Acharya U R. Gabor CNN based intelligent system for visual sentiment analysis of social media data on cloud environment[J]. IEEE Access, 2022, 10: 132455-132471.
[17] Ou H C, Qing C M, Xu X M, et al. Multi-level context pyramid network for visual sentiment analysis[J]. Sensors, 2021, 21(6): 2136.
[18] Yao T, Li Y H, Pan Y W, et al. HIRI-ViT: scaling vision transformer with high resolution inputs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(9): 6431-6442.
[19] Yang Z D, Li Z, Zeng A L, et al. ViTKD: feature-based knowledge distillation for vision transformers[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle, WA, USA: IEEE, 2024: 1379-1388.
[20] Kumar G M K, Mendola J, Shmuel A. Nestedmorph: enhancing deformable medical image registration with nested attention mechanisms[C]//2025 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Tucson, AZ, USA: IEEE, 2025: 4683-4692.
[21] Zhu T, Li L D, Yang J F, et al. Multimodal sentiment analysis with image-text interaction network[J]. IEEE Transactions on Multimedia, 2023, 25: 3375-3385.
[22] Yu J F, Chen K, Xia R. Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis[J]. IEEE Transactions on Affective Computing, 2023, 14(3): 1966-1978.
[23] Bernardi M L, Cimitile M. Report generation from X-ray imaging by retrieval-augmented generation and improved image-text matching[C]//2024 International Joint Conference on Neural Networks(IJCNN). Yokohama, Japan: IEEE, 2024: 1-8.
[24] Kim S, Jo D, Lee D, et al. MAGVLT: masked generative vision-and-language transformer[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, BC, Canada: IEEE, 2023: 23338-23348.
[25] Tsai Y H, Bai S J, Liang P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy. Stroudsburg, PA, USA: ACL, 2019: 6558-6569.
基本信息:
中图分类号:TP391.1;TP391.41
引用信息:
[1]张文波,侯亚迪,王梦煊.面向情感分析的双流时序增强与贡献权重融合模型[J].沈阳理工大学学报().
Citation Information:
[1]ZHANG Wenbo,HOU Yadi,WANG Mengxuan.A Dual-Stream Temporal Enhancement and Contribution Weight Fusion Model for Sentiment Analysis[J].沈阳理工大学学报 Journal of Shenyang Ligong University().
基金信息:
辽宁省科技厅2023年人工智能领域应用基础研究计划项目(2023JH26/10300007)
2026-06-01
2026-06-01
2026-06-01