| [1] |
包翠竹,丁凯,董建峰,等. 视频问答技术研究进展[J]. 计算机研究与发展, 2024, 61(3): 639-673.
|
|
Bao Cuizhu, Ding Kai, Dong Jianfeng, et al. Research progress of video question answering technologies[J]. Journal of Computer Research and Development, 2024, 61(3): 639-673.
|
| [2] |
Zong L, Wan J, Zhang X, et al. Video-context aligned transformer for video question answering[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(17): 19795-19803.
|
| [3] |
Xu H, Ye Q, Yan M, et al. mPLUG-2: A modularized multi-modal foundation model across text, image and video[J]. Proceedings of Machine Learning Research, 2023, 202: 38728-38748.
|
| [4] |
Gao D, Zhou L, Zhu L, et al. MIST: multi-modal iterative spatial-temporal transformer for long-form video question answering[C]// CVPR 2023. Piscataway: IEEE, 2023: 14773-14783.
|
| [5] |
Li Y, Xiao J, Feng C, et al. Discovering spatio-temporal rationales for video question answering[C]// ICCV 2023. Piscataway: IEEE, 2023: 13823-13832.
|
| [6] |
Lai X, Tian Z, Chen Y, et al. LISA: reasoning segmentation via large language model[C]// CVPR 2024. Piscataway: IEEE, 2024: 9579-9589.
|
| [7] |
Zhang Y, Wu J, Li W, et al. LLaVA-Video: video instruction tuning with synthetic data[PP/OL]. V2. arXiv (2025-08-01) [2025-09-10]..
|
| [8] |
Zhu B, Lin B, Ning M, et al. LanguageBind: extending video-language pretraining to n-modality by language-based semantic alignment[PP/OL]. V7. arXiv (2024-01-22) [2025-05-25]..
|
| [9] |
Li K, He Y, Wang Y, et al. VideoChat: chat-centric video understanding[J]. SCIENCE CHINA Information Sciences, 2025, 68(10): No.200102.
|
| [10] |
Xiao J, Shang X, Yao A, et al. NExT-QA: next phase of question-answering to explaining temporal actions[C]// CVPR 2021. Piscataway: IEEE, 2021: 9772-9781.
|
| [11] |
Mangalam K, Akshkulakov R, Malik J. EgoSchema: a diagnostic benchmark for very long-form video language understanding[C]// NeurIPS 2023. Red Hook: Curran Associates Inc., 2023: 46212-46244.
|
| [12] |
Xu L, Zhao Y, Zhou D, et al. PLLaVA: parameter-free LLaVA extension from images to videos for video dense captioning[PP/OL]. V2. arXiv (2024-04-29) [2025-03-20]..
|
| [13] |
Lu H, Salah A A, Poppe R. Snakes and Ladders: two steps up for VideoMamba[C]// ICCV 2025. Piscataway: IEEE, 2025: 24234-24244.
|
| [14] |
Lin B, Ye Y, Zhu B, et al. Video-LLaVA: learning united visual representation by alignment before projection[C]// EMNLP 2024. Stroudsburg: ACL, 2024: 5971-5984.
|
| [15] |
Song E, Chai W, Wang G, et al. MovieChat: from dense token to sparse memory for long video understanding[C]// CVPR 2024. Piscataway: IEEE, 2024: 18221-18232.
|
| [16] |
Wang X, Zhang Y, Zohar O, et al. VideoAgent: long-form video understanding with large language model as agent[C]// ECCV 2024, LNCS 15138. Cham: Springer, 2025: 58-76.
|
| [17] |
Zhang C, Lu T, Islam M M, et al. A simple LLM framework for long-range video question-answering[C]// EMNLP 2024 . Stroudsburg: ACL, 2024: 21715-21737.
|
| [18] |
Fan S, Guo M H, Yang S. Agentic keyframe search for video question answering[PP/OL]. arXiv (2025-03-20) [2025-05-21]..
|
| [19] |
Wang S, Zhao Q, Do M Q, et al. Vamos: versatile action models for video understanding[C]// ECCV 2024, LNCS 15070. Cham: Springer, 2025: 142-160.
|
| [20] |
OpenAI. GPT-4 technical report[PP/OL]. V6. arXiv (2024-03-04) [2025-04-20]..
|
| [21] |
Yu S, Cho J, Yadav P, et al. Self-chained image-language model for video localization and question answering[C]// NeurIPS 2023. Red Hook: Curran Associates Inc., 2023: 76749-76771.
|
| [22] |
Kahatapitiya K, Ranasinghe K, Park J, et al. Language repository for long video understanding[C]// Findings of the Association for Computational Linguistics: ACL 2025. Stroudsburg: ACL, 2025: 5627-5646.
|
| [23] |
Min J, Buch S, Nagrani A, et al. MoReVQA: exploring modular reasoning models for video question answering[C]// CVPR 2024. Piscataway: IEEE, 2024: 13235-13245.
|
| [24] |
Wang Z, Yu S, Stengel-Eskin E, et al. VideoTree: adaptive tree-based video representation for LLM reasoning on long videos[C]// CVPR 2025. Piscataway: IEEE, 2025: 3272-3283.
|